Ya Zhang 0002

dblp:85/3714-2 · DBLP profile ↗
← Back
235ranked-venue papers
8as first author
150since 2021 · last 2026
0000-0002-5390-9053ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 143 · 6 first-author · 93 since 2021Graphics, computer vision, multimedia, augmented reality and games · 120 · 79 since 2021Applied, interdisciplinary, general and emerging computing · 32 · 2 first-author · 24 since 2021Databases, data management, data science and information retrieval · 30 · 1 first-author · 5 since 2021Computer networks · 1Theory of computation · 1
YearPublicationVenuePosition
2026 SceneGen: Single-Image 3D Scene Generation in One Feedforward Pass
abstract
3D content generation has recently attracted significant research interest, driven by its critical applications in VR/AR and embodied AI. In this work, we tackle the challenging task of synthesizing multiple 3D assets within a single scene image. Concretely, our contributions are fourfold: (i) we present SceneGen, a novel framework that takes a scene image and corresponding object masks as input, simultaneously producing multiple 3D assets with geometry and texture. Notably, SceneGen operates with no need for extra optimization or asset retrieval; (ii) we introduce a novel feature aggregation module that integrates local and global scene information from visual and geometric encoders within the feature extraction module. Coupled with a position head, this enables the generation of 3D assets and their relative spatial positions in a single feedforward pass; (iii) we demonstrate SceneGen's direct extensibility to multi-image input scenarios. Despite being trained solely on single-image inputs, our architecture yields improved generation performance when multiple images are provided; and (iv) extensive quantitative and qualitative evaluations confirm the efficiency and robustness of our approach. We believe this paradigm offers a novel solution for high-quality 3D content generation, potentially advancing its practical applications in downstream tasks. The code and model will be publicly available at: https://mengmouxu.github.io/SceneGen. “Everything you can imagine is real.” —Pablo Picasso
Yanxu Meng, Haoning Wu 0002, Ya Zhang 0002, Weidi Xie
3DV3
2026 MedS³: Towards Medical Slow Thinking with Self-Evolved Soft Dual-sided Process Supervision
abstract
Medical language models face critical barriers to real-world clinical reasoning applications. However, mainstream efforts, which fall short in task coverage, lack fine-grained supervision for intermediate reasoning steps, and rely on proprietary systems, are still far from a versatile, credible and efficient language model for clinical reasoning usage. To this end, we propose MedS3, a self-evolving framework that imparts robust reasoning capabilities to small, deployable models. Starting with 8,000 curated instances sampled via a curriculum strategy across five medical domains and 16 datasets, we use a small base policy model to conduct Monte Carlo Tree Search (MCTS) for constructing rule-verifiable reasoning trajectories. Self-explored reasoning trajectories ranked by node values are used to bootstrap the policy model via reinforcement fine-tuning and preference learning. Moreover, we introduce a soft dual process reward model that incorporates value dynamics: steps that degrade node value are penalized, enabling fine-grained identification of reasoning errors even when the final answer is correct. Experiments on eleven benchmarks show that MedS3 outperforms the previous state-of-the-art medical model by +6.45 accuracy points and surpasses 32B-scale general-purpose reasoning models by +8.57 points. Additional empirical analysis further demonstrates that MedS3 achieves robust and faithful reasoning behavior.
Shuyang Jiang, Yusheng Liao, Zhe Chen 0024, Ya Zhang 0002, Yanfeng Wang 0001, Yu Wang 0027
AAAI4
2026 Versatile Vision-Language Model for 3D Computed Tomography
abstract
Representation learning serves as a foundational component of medical vision-language models (MVLMs), enabling cross-modal alignment, semantic consistency, and enhanced generalization capabilities for downstream tasks. As generalist models rapidly evolve, there is a pressing need to unify diverse downstream tasks, such as diagnosis, segmentation, report generation, and multiple choice within a cohesive framework, demanding more efficient and versatile visual representation learning. However, current MVLMs predominately follow CLIP-style vision pretraining, failing to leverage heterogeneous data resources with multi-dimensional imaging and diverse annotation forms. And there lacks systematic analysis of efficient vision encoder design across varied downstream applications, including diagnosis, segmentation, and text generation tasks, particularly for volumetric imaging like Computed Tomography (CT). Besides, current MVLMs exhibit constrained voxel-level capabilities, lacking effective multi-task instruction tuning framework capable of achieving robust performance across various downstream tasks. To address these challenges, we propose CTInstruct, a novel MVLM employing a hybrid ResNet-ViT encoder with multi-granular vision-language pretraining for efficient heterogeneous data modeling, and unified instruction tuning that jointly optimizes discriminative, generative, and voxel-level reasoning for volumetric medical imaging. CTInstruct achieves SOTA performance across 8 CT benchmarks, setting a new standard for data-efficient multimodal learning in medical imaging.
Jiayu Lei, Ziqing Fan, Yanyong Zhang, Weidi Xie, Ya Zhang 0002, Yanfeng Wang 0001
AAAI5
2026 Miner: Mining Intrinsic Mastery for Data-Efficient RL in Large Reasoning Models
abstract
Current critic-free RL methods for large reasoning models suffer from severe inefficiency when training on positive homogeneous prompts (where all rollouts are correct), resulting in waste of rollouts due to zero advantage estimates.We introduce a radically simple yet powerful solution to Mine intrinsic mastery (MINER), that repurposes the policy's intrinsic uncertainty as a self-supervised reward signal, with no external supervision, auxiliary models, or additional inference cost.Our method pioneers two key innovations: (1) a token-level focal credit assignment mechanism that dynamically amplifies gradients on critical uncertain tokens while suppressing overconfident ones, and (2) adaptive advantage calibration to seamlessly integrate intrinsic and verifiable rewards.Evaluated across six reasoning benchmarks on Qwen3-4B and Qwen3-8B base models, MINER achieves state-of-theart performance among the other four algorithms, yielding up to 4.58 absolute gains in Pass@1 and 6.66 gains in Pass@K compared to GRPO.Comparison with other methods targeted at exploration enhancement further discloses the superiority of the two newly proposed innovations.This demonstrates that latent uncertainty exploitation is both necessary and sufficient for efficient and scalable RL training of reasoning models.Code is available at https://github.com/pixas/Miner.
Shuyang Jiang, Ya Zhang 0002, Yanfeng Wang 0001, Yu Wang 0027
ACL (1)3
2026 Prompt tuning with preference ranking for few-shot pre-trained decision transformer
Shengchao Hu, Li Shen 0008, Ya Zhang 0002, Dacheng Tao
Sci. China Inf. Sci.3
2026 Genetic-Enhanced Cross-Entropy reinforcement learning
Ya Zhang 0002, Zheng Fu
Neurocomputing1
2026 Privileged information assisted learning from noisy correspondence
Zihua Zhao, Tianjie Dai, Mengxi Chen, Jiangchao Yao, Bo Han 0003, Ya Zhang 0002, Yanfeng Wang 0001
Neurocomputing6
2026 Dual-granularity Sinkhorn Distillation for Enhanced Learning from Long-Tailed Noisy Data
Feng Hong 0004, Zihua Zhao, Zhihan Zhou 0002, Jiangchao Yao, Dongsheng Li 0002, Ya Zhang 0002, Yanfeng Wang 0001
Mach. Learn.7
2026 USGS: Enhancing sparse view synthesis with unseen viewpoint regularization in 3D Gaussian splatting
Ya Zhang 0002, Jiangshu Wei
Pattern Recognit.1
2026 Interpretable Brain MRI Report Generation Anchored by Lesion Topography
abstract
Radiologists face increasing workloads that make accurate and timely report generation both critical and challenging. This paper presents a novel system for grounded automatic brain MRI report generation, with contributions in three key areas: First, we release RadGenome-Brain MRI, a benchmark dataset featuring multi-modal scans, expert-annotated abnormality masks, and radiology reports with region-level grounding to support fine-grained, explainable report generation. Second, we propose AutoRG-Brain, the first brain MRI report generation framework that combines automatic anomaly segmentation with a visual prompting-based language model to produce structured, anatomically grounded findings. Third, we conduct extensive quantitative and expert evaluations across segmentation and reporting tasks, and demonstrate in real clinical settings that our system significantly enhances junior radiologists' ability to detect subtle abnormalities and compose high-quality reports, narrowing the gap with senior doctors. All code, models, and datasets will be publicly released to facilitate future research and development.
Jiayu Lei, Xiaoman Zhang, Chaoyi Wu, Lisong Dai, Ya Zhang 0002, Yanyong Zhang, Yanfeng Wang 0001, Weidi Xie
IEEE J. Biomed. Health Informatics5
2026 Communication Learning in Multi-Agent Systems From Graph Modeling Perspective
abstract
In numerous artificial intelligence applications, the collaborative efforts of multiple intelligent agents are imperative for the successful attainment of target objectives. To enhance coordination among these agents, a distributed communication framework is often employed, wherein each agent must be capable of encoding information received from the environment and determining how to share it with other agents as required by the task at hand. However, indiscriminate information sharing among all agents can be resource-intensive, and the adoption of manually pre-defined communication architectures imposes constraints on inter-agent communication, thus limiting the potential for effective collaboration. Moreover, the communication framework often remains static during inference, which may result in sustained high resource consumption, as in most cases, only key decisions necessitate information sharing among agents. In this study, we propose a novel approach where the communication structure between agents is represented as a learnable graph.We frame this challenge as the task of identifying the optimal communication graph while allowing the architecture parameters to be updated through regular optimization, which requires a bi-level optimization process. By applying continuous relaxation to the graph structure and integrating attention mechanisms, our method, CommFormer, effectively optimizes the communication graph and simultaneously refines the architectural parameters via gradient descent in an end-to-end manner. Additionally, we introduce a temporal gating mechanism for each agent, enabling dynamic decisions on whether to receive shared information at a given time, based on current observations, thus improving decisionmaking efficiency. Comprehensive experiments conducted across a range of cooperative tasks demonstrate the robustness of our model. Our approach enables agents to develop more coordinated and sophisticated strategies, maintaining effectiveness even with varying agent counts.
Shengchao Hu, Ziqing Fan, Li Shen 0008, Ya Zhang 0002, Dacheng Tao
IEEE Trans. Knowl. Data Eng.4
2026 Positional Prompts-Enhanced Brain-Heart-Gut Interactions for Mild Cognitive Impairment Diagnosis
abstract
Mild cognitive impairment (MCI) is the prodromal stage of dementia involving complex interactions between the brain and peripheral organs. Emerging evidence indicates that heart dysfunction and gut microbiota dysbiosis can contribute to MCI pathogenesis. Yet, these discoveries of cross-organ interactions have not been applied to assist MCI diagnosis. In this work, we propose a novel diagnostic framework that exploits the interactions of brain, heart, and gut using whole-body PET images to guide MCI diagnosis for scenarios when only brain MRI, PET, or PET&MRI are available. Specifically, we collected a multi-cohort, multi-modal dataset comprising 1,545 whole-body PET images, 6,010 brain MR images, and 2,446 brain PET images from eight data centers. Organ-specific image encoders are first pretrained for the brain, heart, and gut individually. Then, to effectively align and integrate brain, heart, and gut features, we introduce positional prompts to act as anatomical-level attention to highlight disease-relevant spatial regions, and further develop hierarchical Transformers to model brain-heart, brain-gut, and brain-heart-gut interactions. Finally, to achieve MCI diagnosis using only brain images, we transfer the above brain-heart-gut model to a brain-only model via an introduced multi-level knowledge distillation scheme, including sample-level contrastive distillation, group-level distribution alignment, and response-level supervision. Extensive experiments on multi-center data demonstrate the superiority of our method over the state-of-the-art methods by resorting to effective integration of heart and gut interactions for MCI diagnosis.
Shilun Zhao, Shuwei Bai, Dengqiang Jia, Jiangtao Liang, Han Zhang 0002, Ya Zhang 0002, Zhongxiang Ding, Yin Xu 0001, Kaicong Sun, Dinggang Shen
IEEE Trans. Medical Imaging8
2025 AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation
abstract
With the proliferation of large language models (LLMs) in the medical domain, there is increasing demand for improved evaluation techniques to assess their capabilities. However, traditional metrics like F1 and ROUGE, which rely on token overlaps to measure quality, significantly overlook the importance of medical terminology. While human evaluation tends to be more reliable, it can be very costly and may as well suffer from inaccuracies due to limits in human expertise and motivation. Although there are some evaluation methods based on LLMs, their usability in the medical field is limited due to their proprietary nature or lack of expertise. To tackle these challenges, we present AutoMedEval, an open-sourced automatic evaluation model with 13B parameters specifically engineered to measure the question-answering proficiency of medical LLMs. The overarching objective of AutoMedEval is to assess the quality of responses produced by diverse models, aspiring to significantly reduce the dependence on human evaluation. Specifically, we propose a hierarchical training method involving curriculum instruction tuning and an iterative knowledge introspection mechanism, enabling AutoMedEval to acquire professional medical assessment capabilities with limited instructional data. Human evaluations indicate that AutoMedEval surpasses other baselines in terms of correlation with human judgments.
Xiechi Zhang, Zetian Ouyang, Gerard de Melo, Zhu Cao, Xiaoling Wang 0004, Ya Zhang 0002, Yanfeng Wang 0001, Liang He 0001
ACL (1)7
2025 Multi-modal Medical Diagnosis via Large-small Model Collaboration
abstract
Recent advances in medical AI have shown a clear trend towards large models in healthcare. However, developing large models for multi-modal medical diagnosis remains challenging due to a lack of sufficient modal-complete medical data. Most existing multi-modal diagnostic models are relatively small and struggle with limited feature extraction capabilities. To bridge this gap, we propose AdaCoMed, an adaptive collaborative-learning framework that synergistically integrates the off-the-shelf medical single-modal large models with multi-modal small models. Our framework first employs a mixture-of-modality-experts (MoME) architecture to combine features extracted from multiple single-modal medical large models, and then introduces a novel adaptive co-learning mechanism to collaborate with a multi-modal small model. This co-learning mechanism, guided by an adaptive weighting strategy, dynamically balances the complementary strengths between the MoMEfused large model features and the cross-modal reasoning capabilities of the small model. Extensive experiments on two representative multi-modal medical datasets (MIMICIV-MM and MMIST ccRCC) across six modalities and four diagnostic tasks demonstrate consistent improvements over state-of-the-art baselines, making it a promising solution for real-world medical diagnosis applications. The code is available at https://github.com/Zoew420/AdaCoMed.
Zihua Zhao, Jiangchao Yao, Ya Zhang 0002, Jiajun Bu, Haishuai Wang
CVPR4
2025 Towards Universal Soccer Video Understanding
abstract
As a globally celebrated sport, soccer has attracted widespread interest from fans all over the world. This paper aims to develop a comprehensive multi-modal framework for soccer video understanding. Specifically, we make the following contributions in this paper: (i) we introduce SoccerReplay-1988, the largest multi-modal soccer dataset to date, featuring videos and detailed annotations from 1,988 complete matches, with an automated annotation pipeline; (ii) we present an advanced soccer-specific visual encoder, MatchVision, which leverages spatiotemporal information across soccer videos and excels in various downstream tasks; (iii) we conduct extensive experiments and ablation studies on event classification, commentary generation, and multi-view foul recognition. MatchVision demonstrates state-of-the-art performance on all of them, substantially outperforming existing models, which highlights the superiority of our proposed data and model. We believe that this work will offer a standard paradigm for sports understanding research.“Football is one of the world’s best means of communication. It is impartial, apolitical, and universal.”—— Franz Beckenbauer (1945 - 2024)
Jiayuan Rao, Haoning Wu 0002, Ya Zhang 0002, Yanfeng Wang 0001, Weidi Xie
CVPR4
2025 DSVD: Dynamic Self-Verify Decoding for Faithful Generation in Large Language Models
abstract
The reliability of large language models remains a critical challenge, particularly due to their susceptibility to hallucinations and factual inaccuracies during text generation.Existing solutions either underutilize models' selfcorrection with preemptive strategies or use costly post-hoc verification.To further explore the potential of real-time self-verification and correction, we present Dynamic Self-Verify Decoding (DSVD), a novel decoding framework that enhances generation reliability through real-time hallucination detection and efficient error correction.DSVD integrates two key components: (1) parallel self-verification architecture for continuous quality assessment, (2) dynamic rollback mechanism for targeted error recovery.Extensive experiments across five benchmarks demonstrate DSVD's effectiveness, achieving significant improvement in truthfulness (Quesetion-Answering) and factual accuracy (FActScore).Results show the DSVD can be further incorporated with existing faithful decoding methods to achieve stronger performance.Our work establishes that real-time self-verification during generation offers a viable path toward more faithful language models without sacrificing practical deployability.
YiQiu Guo, Zhe Chen 0024, Pingjie Wang, Yusheng Liao, Ya Zhang 0002, Yanfeng Wang 0001, Yu Wang 0027
EMNLP6
2025 FreeSegDiff: Annotation-free Saliency Segmentation with Diffusion Models
abstract
Learning from a large corpus of data, pre-trained models have achieved impressive progress nowadays. As a popular generative pre-training method, diffusion models stand out by capturing both low-level visual knowledge and high-level semantic relations. In this paper, we propose to exploit such knowledgeable pre-trained diffusion models for mainstream discriminative tasks such as annotation-free saliency segmentation. However, a notable structural discrepancy between generative and discriminative models poses a significant challenge to diffusion models’ direct application. Furthermore, the absence of explicit manually labeled data is a substantial barrier in annotation-free settings. To tackle these issues, we introduce FreeSegDiff, one novel synthesis-exploitation framework containing two-stage strategies. In the first synthesis stage, to alleviate data insufficiency, we synthesize abundant images, and propose a novel training-free DiffusionCut to produce masks. In the second exploitation stage, to bridge the structural gap, we employ the inversion technique to convert given images back to diffusion features. These features seamlessly integrate with downstream architectures. Extensive experiments and ablation studies demonstrate the superiority of adapting diffusion for annotation-free saliency segmentation.
Chaofan Ma, Yuhuan Yang, Chen Ju, Ya Zhang 0002, Yanfeng Wang 0001
ICASSP5
2025 Contrast-Unity for Partially-Supervised Temporal Sentence Grounding
abstract
Temporal sentence grounding aims to detect event timestamps described by the natural language query from given untrimmed videos. The existing fully-supervised setting achieves great results but requires expensive annotation costs; while the weakly-supervised setting adopts cheap labels but performs poorly. To pursue high performance with less annotation costs, this paper introduces an intermediate partially-supervised setting, i.e., only short-clip is available during training. To make full use of partial labels, we specially design one contrast-unity framework, with the two-stage goal of implicit-explicit progressive grounding. In the implicit stage, we align event-query representations at fine granularity using comprehensive quadruple contrastive learning: event-query gather, event-background separation, intra-cluster compactness and inter-cluster separability. Then, high-quality representations bring acceptable grounding pseudo-labels. In the explicit stage, to explicitly optimize grounding objectives, we train one fully-supervised model using obtained pseudo-labels for grounding refinement and denoising. Extensive experiments and thoroughly ablations on Charades-STA and ActivityNet Captions demonstrate the significance of partial supervision, as well as our superior performance.
Haicheng Wang, Chen Ju, Weixiong Lin, Chaofan Ma, Ya Zhang 0002, Yanfeng Wang 0001
ICASSP6
2025 AuscMLLM: Bridging Classification and Reasoning in Heart Sound Analysis with a Multimodal Large Language Model
abstract
This study introduces a multimodal large language model capable of not only accomplishing various heart sound tasks but also providing reasoning, marking an advancement in the field of medical diagnostics. The model’s innovation stems from a collaboration with experts to collect a novel dataset designed specifically for reasoning tasks, addressing the limitations of existing datasets that lacked this capability. Our model integrates multiple novel methodologies to enhance diagnostic accuracy, including the incorporation of knowledge from relevant textbooks through pre-training, the employment of an audio feature extractor optimized for heart sound-text alignment, and a logit adjustment loss tailored for large language model to mitigate the challenge of imbalanced data categories. This approach not only sets a new standard for heart sound analysis but also paves the way for more interpretable and comprehensive diagnostic models in healthcare.
Pingjie Wang, Liudan Zhao, Ya Zhang 0002, Xin Sun 0020, Yanfeng Wang 0001, Yu Wang 0027
ICASSP4
2025 MRGen: Segmentation Data Engine for Underrepresented MRI Modalities
Haoning Wu 0002, Ya Zhang 0002, Yanfeng Wang 0001, Weidi Xie
ICCV3
2025 Differential-Informed Sample Selection Accelerates Multimodal Contrastive Learning
abstract
The remarkable success of contrastive-learning-based multimodal models has been greatly driven by training on ever-larger datasets with expensive compute consumption. Sample selection as an alternative efficient paradigm plays an important direction to accelerate the training process. However, recent advances on sample selection either mostly rely on an oracle model to offline select a high-quality coreset, which is limited in the cold-start scenarios, or focus on online selection based on real-time model predictions, which has not sufficiently or efficiently considered the noisy correspondence. To address this dilemma, we propose a novel Differential-Informed Sample Selection (DISSect) method, which accurately and efficiently discriminates the noisy correspondence for training acceleration. Specifically, we rethink the impact of noisy correspondence on contrastive learning and propose that the differential between the predicted correlation of the current model and that of a historical model is more informative to characterize sample quality. Based on this, we construct a robust differential-based sample selection and analyze its theoretical insights. Extensive experiments on three benchmark datasets and various downstream tasks demonstrate the consistent superiority of DISSect over current state-of-the-art methods. Source code is available at: https://github.com/MediaBrain-SJTU/DISSect.
Zihua Zhao, Feng Hong 0004, Mengxi Chen, Pengyi Chen, Benyuan Liu, Jiangchao Yao, Ya Zhang 0002, Yanfeng Wang 0001
ICCV7
2025 Combatting Dimensional Collapse in LLM Pre-Training Data via Submodular File Selection
abstract
Selecting high-quality pre-training data for large language models (LLMs) is crucial for enhancing their overall performance under limited computation budget, improving both training and sample efficiency. Recent advancements in file selection primarily rely on using an existing or trained proxy model to assess the similarity of samples to a target domain, such as high quality sources BookCorpus and Wikipedia. However, upon revisiting these methods, the domain-similarity selection criteria demonstrates a diversity dilemma, i.e. dimensional collapse in the feature space, improving performance on the domain-related tasks but causing severe degradation on generic performance.To prevent collapse and enhance diversity, we propose a DiverSified File selection algorithm (DiSF), which selects the most decorrelated text files in the feature space. We approach this with a classical greedy algorithm to achieve more uniform eigenvalues in the feature covariance matrix of the selected texts, analyzing its approximation to the optimal solution under a formulation of $\gamma$-weakly submodular optimization problem. Empirically, we establish a benchmark and conduct extensive experiments on the TinyLlama architecture with models from 120M to 1.1B parameters. Evaluating across nine tasks from the Harness framework, DiSF demonstrates a significant improvement on overall performance. Specifically, DiSF saves 98.5\% of 590M training files in SlimPajama, outperforming the full-data pre-training within a 50B training budget, and achieving about 1.5x training efficiency and 5x data efficiency. Source code is available at: https://github.com/MediaBrain-SJTU/DiSF.git.
Ziqing Fan, Shengchao Hu, Pingjie Wang, Li Shen 0008, Ya Zhang 0002, Dacheng Tao, Yanfeng Wang 0001
ICLR6
2025 Fine-tuning with Reserved Majority for Noise Reduction
abstract
Parameter-efficient fine-tuning (PEFT) has revolutionized supervised fine-tuning, where LoRA and its variants gain the most popularity due to their low training costs and zero inference latency. However, LoRA tuning not only injects knowledgeable features but also noisy hallucination during fine-tuning, which hinders the utilization of tunable parameters with the increasing LoRA rank. In this work, we first investigate in-depth the redundancies among LoRA parameters with substantial empirical studies. Aiming to resemble the learning capacity of high ranks from the findings, we set up a new fine-tuning framework, \textbf{P}arameter-\textbf{Re}dundant \textbf{F}ine-\textbf{T}uning (\preft), which follows the vanilla LoRA tuning process but is required to reduce redundancies before merging LoRA parameters back to pre-trained models. Based on this framework, we propose \textbf{No}ise reduction with \textbf{R}eserved \textbf{M}ajority~(\norm), which decomposes the LoRA parameters into majority parts and redundant parts with random singular value decomposition. The major components are determined by the proposed \search method, specifically employing subspace similarity to confirm the parameter groups that share the highest similarity with the base weight. By employing \norm, we enhance both the learning capacity and benefits from larger ranks, which consistently outperforms both LoRA and other \preft-based methods on various downstream tasks, such as general instruction tuning, math reasoning and code generation. Code is available at \url{https://github.com/pixas/NoRM}.
Shuyang Jiang, Yusheng Liao, Ya Zhang 0002, Yanfeng Wang 0001, Yu Wang 0027
ICLR3
2025 G4Seg: Generation for Inexact Segmentation Refinement with Diffusion Models
abstract
This paper considers the problem of utilizing a large-scale text-to-image diffusion model to tackle the challenging Inexact Segmentation (IS) task. Unlike traditional approaches that rely heavily on discriminative-model-based paradigms or dense visual representations derived from internal attention mechanisms, our method focuses on the intrinsic generative priors in Stable Diffusion (SD). Specifically, we exploit the pattern discrepancies between original images and mask-conditional generated images to facilitate a coarse-to-fine segmentation refinement by establishing a semantic correspondence alignment and updating the foreground probability. Comprehensive quantitative and qualitative experiments validate the effectiveness and superiority of our plug-and-play design, underscoring the potential of leveraging generation discrepancies to model dense representations and encouraging further exploration of generative approaches for solving discriminative tasks.
Fei Zhang 0016, Jiangchao Yao, Ya Zhang 0002, Yanfeng Wang 0001
ICME4
2025 ConText: Driving In-context Learning for Text Removal and Segmentation
abstract
This paper presents the first study on adapting the visual in-context learning (V-ICL) paradigm to optical character recognition tasks, specifically focusing on text removal and segmentation. Most existing V-ICL generalists employ a reasoning-as-reconstruction approach: they turn to using a straightforward image-label compositor as the prompt and query input, and then masking the query label to generate the desired output. This direct prompt confines the model to a challenging single-step reasoning process. To address this, we propose a task-chaining compositor in the form of image-removal-segmentation, providing an enhanced prompt that elicits reasoning with enriched intermediates. Additionally, we introduce context-aware aggregation, integrating the chained prompt pattern into the latent query representation, thereby strengthening the model’s in-context reasoning. We also consider the issue of visual heterogeneity, which complicates the selection of homogeneous demonstrations in text recognition. Accordingly, this is effectively addressed through a simple self-prompting strategy, preventing the model’s in-context learnability from devolving into specialist-like, context-free inference. Collectively, these insights culminate in our ConText model, which achieves new state-of-the-art across both in- and out-of-domain benchmarks. The code is available at https://github.com/Ferenas/ConText.
Fei Zhang 0016, Pei Zhang 0011, Baosong Yang, Fei Huang 0002, Yanfeng Wang 0001, Ya Zhang 0002
ICML6
2025 MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition
abstract
Video understanding is a complex challenge that requires effective modeling of spatial-temporal dynamics. With the success of image foundation models (IFMs) in image understanding, recent approaches have explored parameter-efficient fine-tuning (PEFT) to adapt IFMs for video. However, most of these methods tend to process spatial and temporal information separately, which may fail to capture the full intricacy of video dynamics. In this paper, we propose MoMa, an efficient adapter framework that achieves full spatial-temporal modeling by integrating Mamba's selective state space modeling into IFMs. We propose a novel SeqMod operation to inject spatial-temporal information into pre-trained IFMs, without disrupting their original features. By incorporating SeqMod into a Divide-and-Modulate architecture, MoMa enhances video understanding while maintaining computational efficiency. Extensive experiments on multiple video benchmarks demonstrate the effectiveness of MoMa, achieving superior performance with reduced computational cost. Codes will be released upon publication.
Yuhuan Yang, Chaofan Ma, Zhenjie Mao, Jiangchao Yao, Ya Zhang 0002, Yanfeng Wang 0001
ICML5
2025 Prompt Tuning with Diffusion for Few-Shot Pre-trained Policy Generalization
Shengchao Hu, Wanru Zhao, Weixiong Lin, Li Shen 0008, Ya Zhang 0002, Dacheng Tao
AAMAS5
2025 Brain-Heart-Gut Guided Multi-constraint Knowledge Distillation for Early Alzheimer's Disease Diagnosis
Shilun Zhao, Shuwei Bai, Kai Zhang 0039, Yin Xu 0001, Ya Zhang 0002, Kaicong Sun, Dinggang Shen
MICCAI (15)7
2025 RadIR: A Scalable Framework for Multi-grained Medical Image Retrieval via Radiology Report Mining
Chaoyi Wu, Xiao Zhou 0004, Ya Zhang 0002, Yanfeng Wang 0001, Weidi Xie
MICCAI (5)5
2025 Multi-Agent System for Comprehensive Soccer Understanding
Jiayuan Rao, Haoning Wu 0002, Ya Zhang 0002, Yanfeng Wang 0001, Weidi Xie
ACM Multimedia4
2025 RAD: Towards Trustworthy Retrieval-Augmented Multi-modal Clinical Diagnosis
abstract
Clinical diagnosis is a highly specialized discipline requiring both domain expertise and strict adherence to rigorous guidelines. While current AI-driven medical research predominantly focuses on knowledge graphs or natural text pretraining paradigms to incorporate medical knowledge, these approaches primarily rely on implicitly encoded knowledge within model parameters, neglecting task-specific knowledge required by diverse downstream tasks. To address this limitation, we propose **R**etrieval-**A**ugmented **D**iagnosis (RAD), a novel framework that explicitly injects external knowledge into multimodal models directly on downstream tasks. Specifically, RAD operates through three key mechanisms: retrieval and refinement of disease-centered knowledge from multiple medical sources, a guideline-enhanced contrastive loss that constrains the latent distance between multi-modal features and guideline knowledge, and the dual transformer decoder that employs guidelines as queries to steer cross-modal fusion, aligning the models with clinical diagnostic workflows from guideline acquisition to feature extraction and decision-making. Moreover, recognizing the lack of quantitative evaluation of interpretability for multimodal diagnostic models, we introduce a set of criteria to assess the interpretability from both image and text perspectives. Extensive evaluations across four datasets with different anatomies demonstrate RAD's generalizability, achieving state-of-the-art performance. Furthermore, RAD enables the model to concentrate more precisely on abnormal regions and critical indicators, ensuring evidence-based, trustworthy diagnosis. Our code is available at https://github.com/tdlhl/RAD.
Haolin Li 0001, Tianjie Dai, Zhe Chen 0024, Jiangchao Yao, Ya Zhang 0002, Yanfeng Wang 0001
NeurIPS6
2025 SaFiRe: Saccade-Fixation Reiteration with Mamba for Referring Image Segmentation
abstract
Referring Image Segmentation (RIS) aims to segment the target object in an image given a natural language expression. While recent methods leverage pre-trained vision backbones and more training corpus to achieve impressive results, they predominantly focus on simple expressions—short, clear noun phrases like “red car” or “left girl”. This simplification often reduces RIS to a key word/concept matching problem, limiting the model’s ability to handle referential ambiguity in expressions. In this work, we identify two challenging real-world scenarios: object-distracting expressions, which involve multiple entities with contextual cues, and category-implicit expressions, where the object class is not explicitly stated. To address the challenges, we propose a novel framework, SaFiRe, which mimics the human two-phase cognitive process—first forming a global understanding, then refining it through detail-oriented inspection. This is naturally supported by Mamba’s scan-then-update property, which aligns with our phased design and enables efficient multi-cycle refinement with linear complexity. We further introduce aRefCOCO, a new benchmark designed to evaluate RIS models under ambiguous referring expressions. Extensive experiments on both standard and proposed datasets demonstrate the superiority of SaFiRe over state-of-the-art baselines. Project page: https://zhenjiemao.github.io/SaFiRe/.
Zhenjie Mao, Yuhuan Yang, Chaofan Ma, Dongsheng Jiang, Jiangchao Yao, Ya Zhang 0002, Yanfeng Wang 0001
NeurIPS6
2025 Learning to Instruct for Visual Instruction Tuning
abstract
We propose L2T, an advancement of visual instruction tuning (VIT). While VIT equips Multimodal LLMs (MLLMs) with promising multimodal capabilities, the current design choices for VIT often result in overfitting and shortcut learning, potentially degrading performance. This gap arises from an overemphasis on instruction-following abilities, while neglecting the proactive understanding of visual information. Inspired by this, L2T adopts a simple yet effective approach by incorporating the loss function into both the instruction and response sequences. It seamlessly expands the training data, and regularizes the MLLMs from overly relying on language priors. Based on this merit, L2T achieves a significant relative improvement of up to 9% on comprehensive multimodal benchmarks, requiring no additional training data and incurring negligible computational overhead. Surprisingly, L2T attains exceptional fundamental visual capabilities, yielding up to an 18% improvement in captioning performance, while simultaneously alleviating hallucination in MLLMs. Github code: https://github.com/Feng-Hong/L2T.
Zhihan Zhou 0002, Feng Hong 0004, Jiaan Luo, Yushi Ye, Jiangchao Yao, Dongsheng Li 0002, Bo Han 0003, Ya Zhang 0002, Yanfeng Wang 0001
NeurIPS8
2025 MegaFusion: Extend Diffusion Models towards Higher-resolution Image Generation without Further Tuning
abstract
Diffusion models have emerged as frontrunners in text-to-image generation, but their fixed image resolution during training often leads to challenges in high-resolution image generation, such as semantic deviations and object replication. This paper introduces MegaFusion, a novel approach that extends existing diffusion-based text-to-image models towards efficient higher-resolution generation without additional fine-tuning or adaptation. Specifically, we employ an innovative truncate and relay strategy to bridge the denoising processes across different resolutions, allowing for high-resolution image generation in a coarse-to-fine manner. Moreover, by integrating dilated convolutions and noise re-scheduling, we further adapt the model's priors for higher resolution. The versatility and efficacy of MegaFusion make it universally applicable to both latent-space and pixel-space diffusion models, along with other derivative models. Extensive experiments confirm that MegaFusion significantly boosts the capability of existing models to pro-duce images of megapixels and various aspect ratios, while only requiring about 40% of the original computational cost. Code is available at https://haoningwu3639.github.io/MegaFusion/.
Haoning Wu 0002, Shaocheng Shen, Qiang Hu 0003, Xiaoyun Zhang 0001, Ya Zhang 0002, Yanfeng Wang 0001
WACV5
2025 Graph decision transformer for offline reinforcement learning
Shengchao Hu, Li Shen 0008, Ya Zhang 0002, Dacheng Tao
Sci. China Inf. Sci.3
2025 SLIDE: A Unified Mesh and Texture Generation Framework with Enhanced Geometric Control and Multi-view Consistency
Jinyi Wang, Zhaoyang Lyu, Ben Fei, Jiangchao Yao, Ya Zhang 0002, Bo Dai 0002, Dahua Lin, Ying He 0001, Yanfeng Wang 0001
Int. J. Comput. Vis.5
2025 Fairness-guided federated training for generalization and personalization in cross-silo federated learning
abstract
Cross-silo federated learning (FL), which benefits from relatively abundant data and rich computing power, is drawing increasing focus due to the significant transformations that foundation models (FMs) are instigating in the artificial intelligence field. The intensified data heterogeneity issue of this area, unlike that in cross-device FL, is caused mainly by substantial data volumes and distribution shifts across clients, which requires algorithms to comprehensively consider the personalization and generalization balance. In this paper, we aim to address the objective of generalized and personalized federated learning (GPFL) by enhancing the global model’s cross-domain generalization capabilities and simultaneously improving the personalization performance of local training clients. By investigating the fairness of performance distribution within the federation system, we explore a new connection between generalization gap and aggregation weights established in previous studies, culminating in the fairness-guided federated training for generalization and personalization (FFT-GP) approach. FFT-GP integrates a fairness-aware aggregation (FAA) approach to minimize the generalization gap variance among training clients and a meta-learning strategy that aligns local training with the global model’s feature distribution, thereby balancing generalization and personalization. Our extensive experimental results demonstrate FFT-GP’s superior efficacy compared to existing models, showcasing its potential to enhance FL systems across a variety of practical scenarios.
Ziqing Fan, Jiangchao Yao, Ya Zhang 0002, Yanfeng Wang 0001
Frontiers Inf. Technol. Electron. Eng.4
2025 Uncover the balanced geometry in long-tailed contrastive language-image pretraining
Zhihan Zhou 0002, Yushi Ye, Feng Hong 0004, Peisen Zhao, Jiangchao Yao, Ya Zhang 0002, Qi Tian 0001, Yanfeng Wang 0001
Mach. Learn.6
2025 Redundancy-Adaptive Multimodal Learning for imperfect data
Mengxi Chen, Jiangchao Yao, Linyu Xing, Yu Wang 0027, Ya Zhang 0002, Yanfeng Wang 0001
Neural Networks5
2025 Decouple Before Align: Visual Disentanglement Enhances Prompt Tuning
abstract
Prompt tuning (PT), as an emerging resource-efficient fine-tuning paradigm, has showcased remarkable effectiveness in improving the task-specific transferability of vision-language models. This paper delves into a previously overlooked information asymmetry issue in PT, where the visual modality mostly conveys more context than the object-oriented textual modality. Correspondingly, coarsely aligning these two modalities could result in the biased attention, driving the model to merely focus on the context area. To address this, we propose DAPT, an effective PT framework based on an intuitive decouple-before-align concept. First, we propose to explicitly decouple the visual modality into the foreground and background representation via exploiting coarse-and-fine visual segmenting cues, and then both of these decoupled patterns are aligned with the original foreground texts and the hand-crafted background classes, thereby symmetrically strengthening the modal alignment. To further enhance the visual concentration, we propose a visual pull-push regularization tailored for the foreground-background patterns, directing the original visual representation towards unbiased attention on the region-of-interest object. We demonstrate the power of architecture-free DAPT through few-shot learning, base-to-novel generalization, and data-efficient learning, all of which yield superior performance across prevailing benchmarks.
Fei Zhang 0016, Tianfei Zhou, Jiangchao Yao, Ya Zhang 0002, Ivor W. Tsang, Yanfeng Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 RECISTSurv: Hybrid Multi-Task Transformer for Hepatocellular Carcinoma Response and Survival Evaluation
abstract
Transarterial Chemoembolization (TACE) is a widely applied alternative treatment for patients with hepatocellular carcinoma who are not eligible for liver resection or transplantation. However, the clinical outcomes after TACE are highly heterogeneous. There remains an urgent need for effective and efficient strategies to accurately assess tumor response and predict long-term outcomes using longitudinal and multi-center datasets. To address this challenge, we here introduce RECISTSurv, a novel response-driven Transformer model that integrates multi-task learning with a response-driven co-attention mechanism to simultaneously perform liver and tumor segmentation, predict tumor response to TACE, and estimate overall survival based on longitudinal Computed Tomography (CT) imaging. The proposed Response-driven Co-attention layer models the interactions between pre-TACE and post-TACE features guided by the treatment response embedding. This design enables the model to capture complex relationships between imaging features, treatment response, and survival outcomes, thereby enhancing both prediction accuracy and interpretability. In a multi-center validation study, RECISTSurv-predicted prognosis has demonstrated superior precision than state-of-the-art methods with C-indexes ranging from 0.595 to 0.780. Furthermore, when integrated with multi-modal data, RECISTSurvhas emerged as an independent prognostic factor in all three validation cohorts, with hazard ratio (HR) ranging from 1.693 to 20.7 (P = 0.001-0.042). Our results highlight the potential of RECISTSurvas a powerful tool for personalized treatment planning and outcome prediction in hepatocellular carcinoma patients undergoing TACE. The experimental code is made publicly available at https://github.com/rushier/RECISTSurv.
Rushi Jiao, Qiuping Liu, Yao Zhang 0010, Bangzheng Pu, Bingsen Xue, Kailan Yang, Xisheng Liu, Jinrong Qu, Cheng Jin 0005, Ya Zhang 0002, Yanfeng Wang 0001
IEEE Trans. Image Process.11
2025 Few-Shot Anomaly Detection via Category-Agnostic Registration Learning
abstract
Most existing anomaly detection (AD) methods require a dedicated model for each category. Such a paradigm, despite its promising results, is computationally expensive and inefficient, thereby failing to meet the requirements for real-world applications. Inspired by how humans detect anomalies, by comparing a query image to known normal ones, this article proposes a novel few-shot AD (FSAD) framework. Using a training set of normal images from various categories, registration, aiming to align normal images of the same categories, is leveraged as the proxy task for self-supervised category-agnostic representation learning. At test time, an image and its corresponding support set, consisting of a few normal images from the same category, are supplied, and anomalies are identified by comparing the registered features of the test image to its corresponding support image features. Such a setup enables the model to generalize to novel test categories. It is, to our best knowledge, the first FSAD method that requires no model fine-tuning for novel categories: enabling a single model to be applied to all categories. Extensive experiments demonstrate the effectiveness of the proposed method. Particularly, it improves the current state-of-the-art (SOTA) for FSAD by 11.3% and 8.3% on the MVTec and MPDD benchmarks, respectively. The source code is available at https://github.com/Haoyan-Guan/CAReg.
Chaoqin Huang, Haoyan Guan, Aofan Jiang, Ya Zhang 0002, Michael W. Spratling, Xinchao Wang, Yanfeng Wang 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 MedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models
abstract
The emergence of various medical large language models (LLMs) in the medical domain has highlighted the need for unified evaluation standards, as manual evaluation of LLMs proves to be time-consuming and labor-intensive. To address this issue, we introduce MedBench, a comprehensive benchmark for the Chinese medical domain, comprising 40,041 questions sourced from authentic examination exercises and medical reports of diverse branches of medicine. In particular, this benchmark is composed of four key components: the Chinese Medical Licensing Examination, the Resident Standardization Training Examination, the Doctor In-Charge Qualification Examination, and real-world clinic cases encompassing examinations, diagnoses, and treatments. MedBench replicates the educational progression and clinical practice experiences of doctors in Mainland China, thereby establish- ing itself as a credible benchmark for assessing the mastery of knowledge and reasoning abilities in medical language learning models. We perform extensive experiments and conduct an in-depth analysis from diverse perspectives, which culminate in the following findings: (1) Chinese medical LLMs underperform on this benchmark, highlighting the need for significant advances in clinical knowledge and diagnostic precision. (2) Several general-domain LLMs surprisingly possess considerable medical knowledge. These findings elucidate both the capabilities and limitations of LLMs within the context of MedBench, with the ultimate goal of aiding the medical research community.
Yan Cai 0020, Gerard de Melo, Ya Zhang 0002, Yanfeng Wang 0001, Liang He 0001
AAAI5
2024 Adapting Visual-Language Models for Generalizable Anomaly Detection in Medical Images
abstract
Recent advancements in large-scale visual-language pre-trained models have led to significant progress in zero/few-shot anomaly detection within natural image domains. However, the substantial domain divergence between natural and medical images limits the effectiveness of these methodologies in medical anomaly detection. This paper introduces a novel lightweight multi-level adaptation and comparison framework to repurpose the CLIP model for medical anomaly detection. Our approach integrates multiple residual adapters into the pre-trained visual encoder, enabling a stepwise enhancement of visual features across different levels. This multi-level adaptation is guided by multi-level, pixel-wise visual-language feature alignment loss functions, which recalibrate the model's focus from object semantics in natural imagery to anomaly identification in medical images. The adapted features exhibit improved generalization across various medical data types, even in zero-shot scenarios where the model encounters unseen medical modalities and anatomical regions during training. Our experiments on medical anomaly detection benchmarks demonstrate that our method significantly surpasses current state-of-the-art models, with an average AUC improvement of 6.24% and 7.33% for anomaly classification, 2.03% and 2.37% for anomaly segmentation, under the zero-shot and few-shot settings, respectively. Source code is available at: https://github.com/MediaBrain-SJTU/MVFA-AD
Chaoqin Huang, Aofan Jiang, Ya Zhang 0002, Xinchao Wang, Yanfeng Wang 0001
CVPR4
2024 Audio-Visual Segmentation via Unlabeled Frame Exploitation
abstract
Audio-visual segmentation (AVS) aims to segment the sounding objects in video frames. Although great progress has been witnessed, we experimentally reveal that current methods reach marginal performance gain within the use of the unlabeled frames, leading to the underutilization issue. To fully explore the potential of the unlabeled frames for AVS, we explicitly divide them into two categories based on their temporal characteristics, i.e., neighboring frame (NF) and distantframe (DF). NFs, temporally adjacent to the labeled frame, often contain rich motion information that assists in the accurate localization of sounding objects. Contrary to NFs, DFs have long temporal distaaces from the labeled frame, which share semantic-similar objects with appearance variations. Considering their unique characteristics, we propose a versatile framework that effectively leverages them to tackle AVS. Specifically, for NFs, we exploit the motion cues as the dynamic guidance to improve the objectness localization. Besides, we exploit the semantic cues in DFs by treating them as valid augmentations to the labeled frames, which are then used to enrich data diversity in a self-training manner. Extensive experimental results demonstrate the versatility and superiority of our method, unleashing the power of the abundant unlabeled frames.
Jinxiang Liu, Fei Zhang 0016, Chen Ju, Ya Zhang 0002, Yanfeng Wang 0001
CVPR5
2024 Mitigating Noisy Correspondence by Geometrical Structure Consistency Learning
abstract
Noisy correspondence that refers to mismatches in cross-modal data pairs, is prevalent on human-annotated or web-crawled datasets. Prior approaches to leverage such data mainly consider the application of uni-modal noisy label learning without amending the impact on both cross-modal and intra-modal geometrical structures in multimodal learning. Actually, we find that both structures are effective to discriminate noisy correspondence through structural differences when being wellestablished. Inspired by this observation, we introduce a Geometrical Structure Consistency (GSC) method to infer the true correspon-dence. Specifically, GSC ensures the preservation of geometrical structures within and between modalities, allowing for the accurate discrimination of noisy samples based on structural differences. Utilizing these inferred true correspondence labels, GSC refines the learning of geometrical structures by filtering out the noisy samples. Experiments across four cross-modal datasets confirm that GSC effectively identifies noisy samples and significantly outperforms the current leading methods. Source code is available at: https://github.com/MediaBrain-SJTU/GSC.
Zihua Zhao, Mengxi Chen, Tianjie Dai, Jiangchao Yao, Bo Han 0003, Ya Zhang 0002, Yanfeng Wang 0001
CVPR6
2024 Low-Rank Knowledge Decomposition for Medical Foundation Models
abstract
The popularity of large-scale pretraining has promoted the development of medical foundation models. However, some studies have shown that although foundation models exhibit strong general feature extraction capabilities, their performance on specific tasks is still inferior to task-specific methods. In this paper, we explore a new perspective called “Knowledge Decomposition” to improve the performance on specific medical tasks, which deconstruct the foundation model into multiple lightweight expert models, each dedicated to a particular task, with the goal of improving specialization while concurrently mitigating resource expenditure. To accomplish the above objective, we design a novel framework named Low-Rank Knowledge De-composition (LoRKD), which explicitly separates graidents by incorporating low-rank expert modules and the efficient knowledge separation convolution. Extensive experimental results demonstrate that the decomposed models perform well in terms of performance and transferability, even surpassing the original foundation models. Source code is available at: https://github.com/MediaBrain-SJTU/LoRKD
Haolin Li 0001, Jiangchao Yao, Ya Zhang 0002, Yanfeng Wang 0001
CVPR5
2024 Multi-sentence Grounding for Long-Term Instructional Video
Qirui Chen, Tengda Han, Ya Zhang 0002, Yanfeng Wang 0001, Weidi Xie
ECCV (56)4
2024 ReMamber: Referring Image Segmentation with Mamba Twister
Yuhuan Yang, Chaofan Ma, Jiangchao Yao, Zhun Zhong, Ya Zhang 0002, Yanfeng Wang 0001
ECCV (10)5
2024 Knowledge-Enhanced Visual-Language Pretraining for Computational Pathology
Xiao Zhou 0007, Xiaoman Zhang, Chaoyi Wu, Ya Zhang 0002, Weidi Xie, Yanfeng Wang 0001
ECCV (52)4
2024 CliMedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models in Clinical Scenarios
abstract
With the proliferation of Large Language Models (LLMs) in diverse domains, there is a particular need for unified evaluation standards in Chinese clinical medical scenarios, where models need to be examined very thoroughly.We present CliMedBench, a comprehensive benchmark with 14 expert-guided core clinical scenarios specifically designed to assess the medical ability of LLMs across 7 pivot dimensions 1 .It comprises 33,735 questions derived from real-world medical reports of top-tier tertiary hospitals and authentic examination exercises.The reliability of this benchmark has been confirmed in several ways.Subsequent experiments with existing LLMs have led to the following findings: (i) Chinese medical LLMs underperform on this benchmark, especially where medical reasoning and factual consistency are vital, underscoring the need for advances in clinical knowledge and diagnostic accuracy.(ii) Several general-domain LLMs demonstrate substantial potential in medical clinics, while the limited input capacity of many medical LLMs hinders their practical use.These findings reveal both the strengths and limitations of LLMs in clinical scenarios and offer critical insights for medical research.
Zetian Ouyang, Yishuai Qiu, Gerard de Melo, Ya Zhang 0002, Yanfeng Wang 0001, Liang He 0001
EMNLP5
2024 RaTEScore: A Metric for Radiology Report Generation
abstract
This paper introduces a novel, entity-aware metric, termed as Radiological Report (Text) Evaluation (RaTEScore), to assess the quality of medical reports generated by AI models.RaTEScore emphasizes crucial medical entities, such as diagnostic outcomes and anatomical details.Moreover, it is robust against medical synonyms and sensitive to negation expressions.Technically, we developed a comprehensive medical NER dataset, RaTE-NER, and trained an NER model specifically for this purpose.This model enables the decomposition of complex radiological reports into constituent medical entities.The metric itself is derived by comparing the similarity of entity embeddings, obtained from a language model, based on their types and relevance to clinical significance.Our evaluations demonstrate that RaTEScore aligns more closely with human preference than existing metrics, validated both on established public benchmarks and our newly proposed RaTE-Eval benchmark.
Weike Zhao, Chaoyi Wu, Xiaoman Zhang, Ya Zhang 0002, Yanfeng Wang 0001, Weidi Xie
EMNLP4
2024 Pre-Post Interaction Learning for Brain Tumor Segmentation with Missing MRI Modalities
abstract
Complete multimodal Magnetic Resonance Imaging (MRI) plays an indispensable role in the task of brain tumor segmentation. However, the issue of missing-modality often arises in clinical practice, leading to a significant decline in the accuracy of segmentation. Current methods exhibit suboptimal performance under such scenarios of severe missing modalities due to either limited synthesis capacity for missing modalities or modality-specific information loss caused by strict alignment in latent space. To address this challenge, we propose a Pre-Post Interaction Learning (PPIL) approach that enhances the model’s robustness under severe missing-modality scenarios while maintaining competitive performance when most of the modalities are available. Specifically, separate branches are introduced for each modality to preserve modality-specific information. Meanwhile, a Pre-Interaction component that takes the concatenation of available modalities as input is introduced to capture more inter-modal correlations. Furthermore, a Post-Interaction component is proposed to perceive the importance of all branches and dynamically combine their information, thus mitigating the information loss from the strict latent feature alignment. We validate the effectiveness of PPIL on two benchmark datasets, BraTS2020 and BraTS2018, demonstrating a significant improvement in performance under severe missing-modality scenarios while preserving competitive performance when most modalities are available.
Linyu Xing, Mengxi Chen, Jiangchao Yao, Ya Zhang 0002, Yanfeng Wang 0001
ICASSP4
2024 JOINTRF: End-To-End Joint Optimization for Dynamic Neural Radiance Field Representation and Compression
abstract
Neural Radiance Field (NeRF) excels in photo-realistically static scenes, inspiring numerous efforts to facilitate volumetric videos. However, rendering dynamic and long-sequence radiance fields remains challenging due to the significant data required to represent volumetric videos. In this paper, we propose a novel end-to-end joint optimization scheme of dynamic NeRF representation and compression, called JointRF, thus achieving significantly improved quality and compression efficiency against the previous methods. Specifically, JointRF employs a compact residual feature grid and a coefficient feature grid to represent the dynamic NeRF. This representation handles large motions without compromising quality while concurrently diminishing temporal redundancy. We also introduce a sequential feature compression subnetwork to further reduce spatial-temporal redundancy. Finally, the representation and compression subnetworks are end-to-end trained combined within the JointRF. Extensive experiments demonstrate that JointRF can achieve superior compression performance across various datasets.
Zihan Zheng, Houqiang Zhong, Qiang Hu 0003, Xiaoyun Zhang 0001, Li Song 0001, Ya Zhang 0002, Yanfeng Wang 0001
ICIP6
2024 On Harmonizing Implicit Subpopulations
abstract
Machine learning algorithms learned from data with skewed distributions usually suffer from poor generalization, especially when minority classes matter as much as, or even more than majority ones. This is more challenging on class-balanced data that has some hidden imbalanced subpopulations, since prevalent techniques mainly conduct class-level calibration and cannot perform subpopulation-level adjustments without subpopulation annotations. Regarding implicit subpopulation imbalance, we reveal that the key to alleviating the detrimental effect lies in effective subpopulation discovery with proper rebalancing. We then propose a novel subpopulation-imbalanced learning method called Scatter and HarmonizE (SHE). Our method is built upon the guiding principle of optimal data partition, which involves assigning data to subpopulations in a manner that maximizes the predictive information from inputs to labels. With theoretical guarantees and empirical evidences, SHE succeeds in identifying the hidden subpopulations and encourages subpopulation-balanced predictions. Extensive experiments on various benchmark datasets show the effectiveness of SHE.
Feng Hong 0004, Jiangchao Yao, Yueming Lyu, Zhihan Zhou 0002, Ivor W. Tsang, Ya Zhang 0002, Yanfeng Wang 0001
ICLR6
2024 Learning Multi-Agent Communication from Graph Modeling Perspective
abstract
In numerous artificial intelligence applications, the collaborative efforts of multiple intelligent agents are imperative for the successful attainment of target objectives. To enhance coordination among these agents, a distributed communication framework is often employed. However, information sharing among all agents proves to be resource-intensive, while the adoption of a manually pre-defined communication architecture imposes limitations on inter-agent communication, thereby constraining the potential for collaborative efforts. In this study, we introduce a novel approach wherein we conceptualize the communication architecture among agents as a learnable graph. We formulate this problem as the task of determining the communication graph while enabling the architecture parameters to update normally, thus necessitating a bi-level optimization process. Utilizing continuous relaxation of the graph representation and incorporating attention units, our proposed approach, CommFormer, efficiently optimizes the communication graph and concurrently refines architectural parameters through gradient descent in an end-to-end manner. Extensive experiments on a variety of cooperative tasks substantiate the robustness of our model across diverse cooperative scenarios, where agents are able to develop more coordinated and sophisticated strategies regardless of changes in the number of agents.
Shengchao Hu, Li Shen 0008, Ya Zhang 0002, Dacheng Tao
ICLR3
2024 Domain-Inspired Sharpness-Aware Minimization Under Domain Shifts
abstract
This paper presents a Domain-Inspired Sharpness-Aware Minimization (DISAM) algorithm for optimization under domain shifts. It is motivated by the inconsistent convergence degree of SAM across different domains, which induces optimization bias towards certain domains and thus impairs the overall convergence. To address this issue, we consider the domain-level convergence consistency in the sharpness estimation to prevent the overwhelming (deficient) perturbations for less (well) optimized domains. Specifically, DISAM introduces the constraint of minimizing variance in the domain loss, which allows the elastic gradient calibration in perturbation generation: when one domain is optimized above the averaging level w.r.t. loss, the gradient perturbation towards that domain will be weakened automatically, and vice versa. Under this mechanism, we theoretically show that DISAM can achieve faster overall convergence and improved generalization in principle when inconsistent convergence emerges. Extensive experiments on various domain generalization benchmarks show the superiority of DISAM over a range of state-of-the-art methods. Furthermore, we show the superior efficiency of DISAM in parameter-efficient fine-tuning combined with the pretraining models. The source code is released at https://github.com/MediaBrain-SJTU/DISAM.
Ziqing Fan, Jiangchao Yao, Ya Zhang 0002, Yanfeng Wang 0001
ICLR4
2024 Long-tailed Diffusion Models with Oriented Calibration
abstract
Diffusion models are acclaimed for generating high-quality and diverse images. However, their performance notably degrades when trained on data with a long-tailed distribution. For long tail diffusion model generation, current works focus on the calibration and enhancement of the tail generation with head-tail knowledge transfer. The transfer process relies on the abundant diversity derived from the head class and, more significantly, the condition capacity of the model prediction. However, the dependency on the conditional model prediction to realize the knowledge transfer might exhibit bias during training, leading to unsatisfactory generation results and lack of robustness. Utilizing a Bayesian framework, we develop a weighted denoising score-matching technique for knowledge transfer directly from head to tail classes. Additionally, we incorporate a gating mechanism in the knowledge transfer process. We provide statistical analysis to validate this methodology, revealing that the effectiveness of such knowledge transfer depends on both label distribution and sample similarity, providing the insight to consider sample similarity when re-balancing the label proportion in training. We extensively evaluate our approach with experiments on multiple benchmark datasets, demonstrating its effectiveness and superior performance compared to existing methods. Code: \url{https://github.com/MediaBrain-SJTU/OC_LT}.
Huangjie Zheng, Jiangchao Yao, Xiangfeng Wang 0001, Mingyuan Zhou, Ya Zhang 0002, Yanfeng Wang 0001
ICLR6
2024 MVTexGen: Synthesising 3D Textures Using Multi-View Diffusion
abstract
We introduce MVTexGen, a novel method for generating textures on 3D geometries using a 2D text-to-image diffusion model. Traditional project-and-inpaint techniques, often result in texture inconsistencies due to uneven diffusion across views. We address this issue by integration of a multi-view prior into the generation process, ensuring synchronous view generation and uniformity in overlapping areas. It combines Multi-View Diffusion models with depth-conditioned diffusion models to create consistent, depth-aware texture maps. Addressing latent space gaps, MVTexGen refines the texture map by increasing view count and fusing denoised views for uniformity. Our extensive experiments on benchmark datasets show MVTexGen’s superiority in generating high-quality, detailed textures, outperforming current state-of-the-art methods.
Jinyi Wang, Fei Ben, Huangjie Zheng, Jiangchao Yao, Ya Zhang 0002, Yanfeng Wang 0001
ICME5
2024 Diversified Batch Selection for Training Acceleration
abstract
The remarkable success of modern machine learning models on large datasets often demands extensive training time and resource consumption. To save cost, a prevalent research line, known as online batch selection, explores selecting informative subsets during the training process. Although recent efforts achieve advancements by measuring the impact of each sample on generalization, their reliance on additional reference models inherently limits their practical applications, when there are no such ideal models available. On the other hand, the vanilla reference-model-free methods involve independently scoring and selecting data in a sample-wise manner, which sacrifices the diversity and induces the redundancy. To tackle this dilemma, we propose Diversified Batch Selection (DivBS), which is reference-model-free and can efficiently select diverse and representative samples. Specifically, we define a novel selection objective that measures the group-wise orthogonalized representativeness to combat the redundancy issue of previous sample-wise criteria, and provide a principled selection-efficient realization. Extensive experiments across various tasks demonstrate the significant superiority of DivBS in the performance-speedup trade-off. The code is publicly available.
Feng Hong 0004, Yueming Lyu, Jiangchao Yao, Ya Zhang 0002, Ivor W. Tsang, Yanfeng Wang 0001
ICML4
2024 Locally Estimated Global Perturbations are Better than Local Perturbations for Federated Sharpness-aware Minimization
abstract
In federated learning (FL), the multi-step update and data heterogeneity among clients often lead to a loss landscape with sharper minima, degenerating the performance of the resulted global model. Prevalent federated approaches incorporate sharpness-aware minimization (SAM) into local training to mitigate this problem. However, the local loss landscapes may not accurately reflect the flatness of global loss landscape in heterogeneous environments; as a result, minimizing local sharpness and calculating perturbations on client data might not align the efficacy of SAM in FL with centralized training. To overcome this challenge, we propose FedLESAM, a novel algorithm that locally estimates the direction of global perturbation on client side as the difference between global models received in the previous active and current rounds. Besides the improved quality, FedLESAM also speed up federated SAM-based approaches since it only performs once backpropagation in each iteration. Theoretically, we prove a slightly tighter bound than its original FedSAM by ensuring consistent perturbation. Empirically, we conduct comprehensive experiments on four federated benchmark datasets under three partition strategies to demonstrate the superior performance and efficiency of FedLESAM.
Ziqing Fan, Shengchao Hu, Jiangchao Yao, Gang Niu 0001, Ya Zhang 0002, Masashi Sugiyama, Yanfeng Wang 0001
ICML5
2024 HarmoDT: Harmony Multi-Task Decision Transformer for Offline Reinforcement Learning
abstract
The purpose of offline multi-task reinforcement learning (MTRL) is to develop a unified policy applicable to diverse tasks without the need for online environmental interaction. Recent advancements approach this through sequence modeling, leveraging the Transformer architecture’s scalability and the benefits of parameter sharing to exploit task similarities. However, variations in task content and complexity pose significant challenges in policy formulation, necessitating judicious parameter sharing and management of conflicting gradients for optimal policy performance. In this work, we introduce the Harmony Multi-Task Decision Transformer (HarmoDT), a novel solution designed to identify an optimal harmony subspace of parameters for each task. We approach this as a bi-level optimization problem, employing a meta-learning framework that leverages gradient-based techniques. The upper level of this framework is dedicated to learning a task-specific mask that delineates the harmony subspace, while the inner level focuses on updating parameters to enhance the overall performance of the unified policy. Empirical evaluations on a series of benchmarks demonstrate the superiority of HarmoDT, verifying the effectiveness of our approach.
Shengchao Hu, Ziqing Fan, Li Shen 0008, Ya Zhang 0002, Yanfeng Wang 0001, Dacheng Tao
ICML4
2024 Q-value Regularized Transformer for Offline Reinforcement Learning
abstract
Recent advancements in offline reinforcement learning (RL) have underscored the capabilities of Conditional Sequence Modeling (CSM), a paradigm that learns the action distribution based on history trajectory and target returns for each state. However, these methods often struggle with stitching together optimal trajectories from sub-optimal ones due to the inconsistency between the sampled returns within individual trajectories and the optimal returns across multiple trajectories. Fortunately, Dynamic Programming (DP) methods offer a solution by leveraging a value function to approximate optimal future returns for each state, while these techniques are prone to unstable learning behaviors, particularly in long-horizon and sparse-reward scenarios. Building upon these insights, we propose the Q-value regularized Transformer (QT), which combines the trajectory modeling ability of the Transformer with the predictability of optimal future returns from DP methods. QT learns an action-value function and integrates a term maximizing action-values into the training loss of CSM, which aims to seek optimal actions that align closely with the behavior policy. Empirical evaluations on D4RL benchmark datasets demonstrate the superiority of QT over traditional DP and CSM methods, highlighting the potential of QT to enhance the state-of-the-art in offline RL.
Shengchao Hu, Ziqing Fan, Chaoqin Huang, Li Shen 0008, Ya Zhang 0002, Yanfeng Wang 0001, Dacheng Tao
ICML5
2024 Exploring Training on Heterogeneous Data with Mixture of Low-rank Adapters
abstract
Training a unified model to take multiple targets into account is a trend towards artificial general intelligence. However, how to efficiently mitigate the training conflicts among heterogeneous data collected from different domains or tasks remains under-explored. In this study, we explore to leverage Mixture of Low-rank Adapters (MoLA) to mitigate conflicts in heterogeneous data training, which requires to jointly train the multiple low-rank adapters and their shared backbone. Specifically, we introduce two variants of MoLA, namely, MoLA-Grad and MoLA-Router, to respectively handle the target-aware and target-agnostic scenarios during inference. The former uses task identifiers to assign personalized low-rank adapters to each task, disentangling task-specific knowledge towards their adapters, thereby mitigating heterogeneity conflicts. The latter uses a novel Task-wise Decorrelation (TwD) loss to intervene the router to learn oriented weight combinations of adapters to homogeneous tasks, achieving similar effects. We conduct comprehensive experiments to verify the superiority of MoLA over previous state-of-the-art methods and present in-depth analysis on its working mechanism. Source code is available at: https://github.com/MediaBrain-SJTU/MoLA
Zihua Zhao, Haolin Li 0001, Jiangchao Yao, Ya Zhang 0002, Yanfeng Wang 0001
ICML6
2024 Reprogramming Distillation for Medical Foundation Models
Haolin Li 0001, Jiangchao Yao, Ya Zhang 0002, Yanfeng Wang 0001
MICCAI (11)5
2024 HPC: Hierarchical Progressive Coding Framework for Volumetric Video
abstract
Volumetric video based on Neural Radiance Field (NeRF) holds vast potential for various 3D applications, but its substantial data volume poses significant challenges for compression and transmission. Current NeRF compression lacks the flexibility to adjust video quality and bitrate within a single model for various network and device capacities. To address these issues, we propose HPC, a novel hierarchical progressive volumetric video coding framework achieving variable bitrate using a single model. Specifically, HPC introduces a hierarchical representation with a multi-resolution residual radiance field to reduce temporal redundancy in long-duration sequences while simultaneously generating various levels of detail. Then, we propose an end-to-end progressive learning approach with a multi-rate-distortion loss function to jointly optimize both hierarchical representation and compression. Our HPC trained only once can realize multiple compression levels, while the current methods need to train multiple fixed-bitrate models for different rate-distortion (RD) tradeoffs. Extensive experiments demonstrate that HPC achieves flexible quality levels with variable bitrate by a single model and exhibits competitive RD performance, even outperforming fixed-bitrate models across various datasets.
Zihan Zheng, Houqiang Zhong, Qiang Hu 0003, Xiaoyun Zhang 0001, Li Song 0001, Ya Zhang 0002, Yanfeng Wang 0001
ACM Multimedia6
2024 Probabilistic Conformal Distillation for Enhancing Missing Modality Robustness
abstract
Multimodal models trained on modality-complete data are plagued with severe performance degradation when encountering modality-missing data. Prevalent cross-modal knowledge distillation-based methods precisely align the representation of modality-missing data and that of its modality-complete counterpart to enhance robustness. However, due to the irreparable information asymmetry, this determinate alignment is too stringent, easily inducing modality-missing features to capture spurious factors erroneously. In this paper, a novel multimodal Probabilistic Conformal Distillation (PCD) method is proposed, which considers the inherent indeterminacy in this alignment. Given a modality-missing input, our goal is to learn the unknown Probability Density Function (PDF) of the mapped variables in the modality-complete space, rather than relying on the brute-force point alignment. Specifically, PCD models the modality-missing feature as a probabilistic distribution, enabling it to satisfy two characteristics of the PDF. One is the extremes of probabilities of modality-complete feature points on the PDF, and the other is the geometric consistency between the modeled distributions and the peak points of different PDFs. Extensive experiments on a range of benchmark datasets demonstrate the superiority of PCD over state-of-the-art methods. Code is available at: https://github.com/mxchen-mc/PCD.
Mengxi Chen, Fei Zhang 0016, Zihua Zhao, Jiangchao Yao, Ya Zhang 0002, Yanfeng Wang 0001
NeurIPS5
2024 TAIA: Large Language Models are Out-of-Distribution Data Learners
abstract
Fine-tuning on task-specific question-answer pairs is a predominant method for enhancing the performance of instruction-tuned large language models (LLMs) on downstream tasks. However, in certain specialized domains, such as healthcare or harmless content generation, it is nearly impossible to obtain a large volume of high-quality data that matches the downstream distribution. To improve the performance of LLMs in data-scarce domains with domain-mismatched data, we re-evaluated the Transformer architecture and discovered that not all parameter updates during fine-tuning contribute positively to downstream performance. Our analysis reveals that within the self-attention and feed-forward networks, only the fine-tuned attention parameters are particularly beneficial when the training set's distribution does not fully align with the test set. Based on this insight, we propose an effective inference-time intervention method: \uline{T}raining \uline{A}ll parameters but \uline{I}nferring with only \uline{A}ttention (TAIA). We empirically validate TAIA using two general instruction-tuning datasets and evaluate it on seven downstream tasks involving math, reasoning, and knowledge understanding across LLMs of different parameter sizes and fine-tuning techniques. Our comprehensive experiments demonstrate that TAIA achieves superior improvements compared to both the fully fine-tuned model and the base model in most scenarios, with significant performance gains. The high tolerance of TAIA to data mismatches makes it resistant to jailbreaking tuning and enhances specialized tasks using general data. Code is available in \url{https://github.com/pixas/TAIA_LLM}.
Shuyang Jiang, Yusheng Liao, Ya Zhang 0002, Yanfeng Wang 0001, Yu Wang 0027
NeurIPS3
2024 Revive Re-weighting in Imbalanced Learning by Density Ratio Estimation
abstract
In deep learning, model performance often deteriorates when trained on highly imbalanced datasets, especially when evaluation metrics require robust generalization across underrepresented classes. To address the challenges posed by imbalanced data distributions, this study introduces a novel method utilizing density ratio estimation for dynamic class weight adjustment, termed as Re-weighting with Density Ratio (RDR). Our method adaptively adjusts the importance of each class during training, mitigates overfitting on dominant classes and enhances model adaptability across diverse datasets. Extensive experiments conducted on various large scale benchmark datasets validate the effectiveness of our method. Results demonstrate substantial improvements in generalization capabilities, particularly under severely imbalanced conditions.
Jiaan Luo, Feng Hong 0004, Jiangchao Yao, Bo Han 0003, Ya Zhang 0002, Yanfeng Wang 0001
NeurIPS5
2024 Annotation-free Audio-Visual Segmentation
abstract
The objective of Audio-Visual Segmentation (AVS) is to localise the sounding objects within visual scenes by accurately predicting pixel-wise segmentation masks. To tackle the task, it involves a comprehensive consideration of both the data and model aspects. In this paper, first, we initiate a novel pipeline for generating artificial data for the AVS task without extra manual annotations. We leverage existing image segmentation and audio datasets and match the image-mask pairs with its corresponding audio samples using category labels in segmentation datasets, that allows us to effortlessly compose (image, audio, mask) triplets for training AVS models. The pipeline is annotation-free and scalable to cover a large number of categories. Additionally, we introduce a lightweight model SAMA-AVS which adapts the pre-trained segment anything model (SAM) to the AVS task. By introducing only a small number of trainable parameters with adapters, the proposed model can effectively achieve adequate audio-visual fusion and interaction in the encoding stage with vast majority of parameters fixed. We conduct extensive experiments, and the results show our proposed model remarkably surpasses other competing methods. Moreover, by using the proposed model pretrained with our synthetic data, the performance on real AVSBench data is further improved, achieving 83.17 mIoU on S4 subset and 66.95 mIoU on MS3 set. The project page is https://jinxiang-liu.github.io/anno-free-AVS/.
Jinxiang Liu, Yu Wang 0027, Chen Ju, Chaofan Ma, Ya Zhang 0002, Weidi Xie
WACV5
2024 Multi-modal Prototypes for Open-World Semantic Segmentation
Yuhuan Yang, Chaofan Ma, Chen Ju, Fei Zhang 0016, Jiangchao Yao, Ya Zhang 0002, Yanfeng Wang 0001
Int. J. Comput. Vis.6
2024 PMC-LLaMA: toward building open-source language models for medicine
abstract
OBJECTIVE: Recently, large language models (LLMs) have showcased remarkable capabilities in natural language understanding. While demonstrating proficiency in everyday conversations and question-answering (QA) situations, these models frequently struggle in domains that require precision, such as medical applications, due to their lack of domain-specific knowledge. In this article, we describe the procedure for building a powerful, open-source language model specifically designed for medicine applications, termed as PMC-LLaMA. MATERIALS AND METHODS: We adapt a general-purpose LLM toward the medical domain, involving data-centric knowledge injection through the integration of 4.8M biomedical academic papers and 30K medical textbooks, as well as comprehensive domain-specific instruction fine-tuning, encompassing medical QA, rationale for reasoning, and conversational dialogues with 202M tokens. RESULTS: While evaluating various public medical QA benchmarks and manual rating, our lightweight PMC-LLaMA, which consists of only 13B parameters, exhibits superior performance, even surpassing ChatGPT. All models, codes, and datasets for instruction tuning will be released to the research community. DISCUSSION: Our contributions are 3-fold: (1) we build up an open-source LLM toward the medical domain. We believe the proposed PMC-LLaMA model can promote further development of foundation models in medicine, serving as a medical trainable basic generative language backbone; (2) we conduct thorough ablation studies to demonstrate the effectiveness of each proposed component, demonstrating how different training data and model scales affect medical LLMs; (3) we contribute a large-scale, comprehensive dataset for instruction tuning. CONCLUSION: In this article, we systematically investigate the process of building up an open-source medical-specific LLM, PMC-LLaMA.
Chaoyi Wu, Weixiong Lin, Xiaoman Zhang, Ya Zhang 0002, Weidi Xie, Yanfeng Wang 0001
J. Am. Medical Informatics Assoc.4
2024 Fair evaluation of federated learning algorithms for automated breast density classification: The results of the 2022 ACR-NCI-NVIDIA federated learning challenge
Kendall Schmidt, Ben Bearce, Ken Chang, Laura Coombs, Keyvan Farahani, Marawan Elbatel, Kaouther Mouheb, Robert Martí, Ya Zhang 0002, Yanfeng Wang 0001, Yaojun Hu, Haochao Ying, Yuyang Xu, Conrad Testagrose, Mutlu Demirer, Vikash Gupta, Ünal Akünal, Markus Bujotzek, Klaus H. Maier-Hein, Yi Qin 0006, Xiaomeng Li 0001, Jayashree Kalpathy-Cramer, Holger Roth
Medical Image Anal.10
2024 Dynamic-group-aware networks for multi-agent trajectory prediction with relational reasoning
Chenxin Xu, Yuxi Wei, Bohan Tang, Sheng Yin, Ya Zhang 0002, Siheng Chen, Yanfeng Wang 0001
Neural Networks5
2024 On Transforming Reinforcement Learning With Transformers: The Development Trajectory
abstract
Transformers, originally devised for natural language processing (NLP), have also produced significant successes in computer vision (CV). Due to their strong expression power, researchers are investigating ways to deploy transformers for reinforcement learning (RL), and transformer-based models have manifested their potential in representative RL benchmarks. In this paper, we collect and dissect recent advances concerning the transformation of RL with transformers (transformer-based RL (TRL)) to explore the development trajectory and future trends of this field. We group the existing developments into two categories: architecture enhancements and trajectory optimizations, and examine the main applications of TRL in robotic manipulation, text-based games (TBGs), navigation, and autonomous driving. Architecture enhancement methods consider how to apply the powerful transformer structure to RL problems under the traditional RL framework, facilitating more precise modeling of agents and environments compared to traditional deep RL techniques. However, these methods are still limited by the inherent defects of traditional RL algorithms, such as bootstrapping and the "deadly triad". Trajectory optimization methods treat RL problems as sequence modeling problems and train a joint state-action model over entire trajectories under the behavior cloning framework; such approaches are able to extract policies from static datasets and fully use the long-sequence modeling capabilities of transformers. Given these advancements, the limitations and challenges in TRL are reviewed and proposals regarding future research directions are discussed. We hope that this survey can provide a detailed introduction to TRL and motivate future research in this rapidly developing field.
Shengchao Hu, Li Shen 0008, Ya Zhang 0002, Yixin Chen 0001, Dacheng Tao
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Learning node representations against perturbations
Xu Chen 0026, Yuangang Pan, Ivor W. Tsang, Ya Zhang 0002
Pattern Recognit.4
2024 Balanced Destruction-Reconstruction Dynamics for Memory-Replay Class Incremental Learning
abstract
Class incremental learning (CIL) aims to incrementally update a trained model with the new classes of samples (plasticity) while retaining previously learned ability (stability). To address the most challenging issue in this goal, i.e., catastrophic forgetting, the mainstream paradigm is memory-replay CIL, which consolidates old knowledge by replaying a small number of old classes of samples saved in the memory. Despite effectiveness, the inherent destruction-reconstruction dynamics in memory-replay CIL are an intrinsic limitation: if the old knowledge is severely destructed, it will be quite hard to reconstruct the lossless counterpart. Our theoretical analysis shows that the destruction of old knowledge can be effectively alleviated by balancing the contribution of samples from the current phase and those saved in the memory. Motivated by this theoretical finding, we propose a novel Balanced Destruction-Reconstruction module (BDR) for memory-replay CIL, which can achieve better knowledge reconstruction by reducing the degree of maximal destruction of old knowledge. Specifically, to achieve a better balance between old knowledge and new classes, the proposed BDR module takes into account two factors: the variance in training status across different classes and the quantity imbalance of samples from the current phase and memory. By dynamically manipulating the gradient during training based on these factors, BDR can effectively alleviate knowledge destruction and improve knowledge reconstruction. Extensive experiments on a range of CIL benchmarks have shown that as a lightweight plug-and-play module, BDR can significantly improve the performance of existing state-of-the-art methods with good generalization. Our code is publicly available here.
Jiangchao Yao, Feng Hong 0004, Ya Zhang 0002, Yanfeng Wang 0001
IEEE Trans. Image Process.4
2024 Server-Client Collaborative Distillation for Federated Reinforcement Learning
abstract
Federated Learning (FL) learns a global model in a distributional manner, which does not require local clients to share private data. Such merit has drawn lots of attention in the interaction scenarios, where Federated Reinforcement Learning (FRL) emerges as a cross-field research direction focusing on the robust training of agents. Different from FL, the heterogeneity problem in FRL is more challenging because the data depends on the policy of agents and the environment dynamics. FRL learns to interact under the non-stationary environment feedback, while the typical FL methods aim at handling the constant data heterogeneity. In this article, we are among the first attempts to analyze the heterogeneity problem in FRL and propose an off-policy FRL framework. Specifically, a student–teacher–student model learning and fusion method, termed asServer-Client Collaborative Distillation(SCCD), is introduced. Unlike the traditional FL, we distill all local models on the server side for model fusion. To reduce the variance of the training, a local distillation is also conducted every time the agent receives the global model. Experimentally, we compare SCCD with a range of straightforward combinations between FL methods and RL. The results demonstrate that SCCD has a superior performance in four classical continuous control tasks with non-IID environments.
Weiming Mai, Jiangchao Yao, Chen Gong 0002, Ya Zhang 0002, Yiu-Ming Cheung, Bo Han 0003
ACM Trans. Knowl. Discov. Data4
2024 UniChest: Conquer-and-Divide Pre-Training for Multi-Source Chest X-Ray Classification
abstract
Vision-Language Pre-training (VLP) that utilizes the multi-modal information to promote the training efficiency and effectiveness, has achieved great success in vision recognition of natural domains and shown promise in medical imaging diagnosis for the Chest X-Rays (CXRs). However, current works mainly pay attention to the exploration on single dataset of CXRs, which locks the potential of this powerful paradigm on larger hybrid of multi-source CXRs datasets. We identify that although blending samples from the diverse sources offers the advantages to improve the model generalization, it is still challenging to maintain the consistent superiority for the task of each source due to the existing heterogeneity among sources. To handle this dilemma, we design a Conquer-and-Divide pre-training framework, termed as UniChest, aiming to make full use of the collaboration benefit of multiple sources of CXRs while reducing the negative influence of the source heterogeneity. Specially, the "Conquer" stage in UniChest encourages the model to sufficiently capture multi-source common patterns, and the "Divide" stage helps squeeze personalized patterns into different small experts (query networks). We conduct thorough experiments on many benchmarks, e.g., ChestX-ray14, CheXpert, Vindr-CXR, Shenzhen, Open-I and SIIM-ACR Pneumothorax, verifying the effectiveness of UniChest over a range of baselines, and release our codes and pre-training models at https://github.com/Elfenreigen/UniChest.
Tianjie Dai, Feng Hong 0004, Jiangchao Yao, Ya Zhang 0002, Yanfeng Wang 0001
IEEE Trans. Medical Imaging5
2024 Federated Adversarial Domain Hallucination for Privacy-Preserving Domain Generalization
abstract
Domain generalization aims to reduce the vulnerability of deep neural networks in the out-of-domain distribution scenario. With the recent and increasing data privacy concerns, federated domain generalization, where multiple domains are distributed on different local clients, has become an important research problem and brings new challenges for learning domain-invariant information from separated domains. In this paper, we address the problem of federated domain generalization from the perspective of domain hallucination. We propose a novel federated domain hallucination learning framework, with no additional data exchange between clients other than model weights, based on the idea that a domain hallucination with enlarged prediction uncertainty for the global model is more likely to transform the samples into an unseen domain. These types of desired domain hallucinations are achieved by generating samples that maximize the entropy of the global model and minimize the cross-entropy of the local model, where the latter loss is further introduced to maintain the sample semantics. By training the local models with the learned domain hallucinations, the final model is expected to be more robust to unseen domain shifts. We perform extensive experiments on three object classification benchmarks and one medical image segmentation benchmark. The proposed method outperforms state-of-the-art methods on all the benchmarks, demonstrating its effectiveness.
Qinwei Xu, Ya Zhang 0002, Yiyan Wu 0001, Yanfeng Wang 0001
IEEE Trans. Multim.3
2024 Online Multi-Agent Forecasting With Interpretable Collaborative Graph Neural Networks
abstract
This article considers predicting future statuses of multiple agents in an online fashion by exploiting dynamic interactions in the system. We propose a novel collaborative prediction unit (CoPU), which aggregates the predictions from multiple collaborative predictors according to a collaborative graph. Each collaborative predictor is trained to predict the agent status by integrating the impact of another agent. The edge weights of the collaborative graph reflect the importance of each predictor. The collaborative graph is adjusted online by multiplicative update, which can be motivated by minimizing an explicit objective. With this objective, we also conduct regret analysis to indicate that, along with training, our CoPU achieves similar performance with the best individual collaborative predictor in hindsight. This theoretical interpretability distinguishes our method from many other graph networks. To progressively refine predictions, multiple CoPUs are stacked to form a collaborative graph neural network. Extensive experiments are conducted on three tasks: online simulated trajectory prediction, online human motion prediction, and online traffic speed prediction, and our methods outperform state-of-the-art works on the three tasks by 28.6%, 17.4%, and 21.0% on average, respectively; in addition, the proposed CoGNNs have lower average time costs in one online training/testing iteration than most previous methods.
Maosen Li, Siheng Chen, Yanning Shen, Genjia Liu, Ivor W. Tsang, Ya Zhang 0002
IEEE Trans. Neural Networks Learn. Syst.6
2023 Temporally-Extended Prompts Optimization for SAM in Interactive Medical Image Segmentation
abstract
The Segmentation Anything Model (SAM) has recently emerged as a foundation model for addressing image segmentation. Owing to the intrinsic complexity of medical images and the high annotation cost, the medical image segmentation (MIS) community has been encouraged to investigate SAM’s zero-shot capabilities to facilitate automatic annotation. Inspired by the extraordinary accomplishments of the interactive medical image segmentation (IMIS) paradigm, this paper focuses on assessing the potential of SAM’s zero-shot capabilities within the IMIS paradigm to amplify its benefits in the MIS domain. Regrettably, we observe that SAM’s vulnerability to prompt forms (e.g., points, bounding boxes) becomes notably pronounced in IMIS. This leads us to develop a mechanism that adaptively offers suitable prompt forms for human experts. We refer to the mechanism above as temporally-extended prompts optimization (TEPO) and model it as a Markov decision process, solvable through reinforcement learning. Numerical experiments on the standardized benchmark Brats2020 demonstrate that the learned TEPO agent can further enhance SAM’s zero-shot capability in the MIS context.
Chuyun Shen, Wenhao Li 0001, Ya Zhang 0002, Yanfeng Wang 0001, Xiangfeng Wang 0001
BIBM3
2023 Zero-shot Composed Text-Image Retrieval
Jiangchao Yao, Ya Zhang 0002, Yanfeng Wang 0001, Weidi Xie
BMVC3
2023 Boost Video Frame Interpolation via Motion Adaptation
Haoning Wu 0002, Xiaoyun Zhang 0001, Weidi Xie, Ya Zhang 0002, Yanfeng Wang 0001
BMVC4
2023 Enhanced Multimodal Representation Learning with Cross-modal KD
abstract
This paper explores the tasks of leveraging auxiliary modalities which are only available at training to enhance multimodal representation learning through cross-modal Knowledge Distillation (KD). The widely adopted mutual information maximization-based objective leads to a short-cut solution of the weak teacher, i.e., achieving the maximum mutual information by simply making the teacher model as weak as the student model. To prevent such a weak solution, we introduce an additional objective term, i.e., the mutual information between the teacher and the auxiliary modality model. Besides, to narrow down the information gap between the student and teacher, we further propose to minimize the conditional entropy of the teacher given the student. Novel training schemes based on contrastive learning and adversarial learning are designed to optimize the mutual information and the conditional entropy, respectively. Experimental results on three popular multimodal benchmark datasets have shown that the proposed method outperforms a range of state-of-the-art approaches for video recognition, video retrieval and emotion classification.
Mengxi Chen, Linyu Xing, Yu Wang 0027, Ya Zhang 0002
CVPR4
2023 Distilling Vision-Language Pre-Training to Collaborate with Weakly-Supervised Temporal Action Localization
abstract
Weakly-supervised temporal action localization (WTAL) learns to detect and classify action instances with only category labels. Most methods widely adopt the off-the-shelf Classification-Based Pre-training (CBP) to generate video features for action localization. However, the different optimization objectives between classification and localization, make temporally localized results suffer from the serious incomplete issue. To tackle this issue without additional annotations, this paper considers to distill free action knowledge from Vision-Language Pre-training (VLP), as we surprisingly observe that the localization results of vanilla VLP have an over-complete issue, which is just complementary to the CBP results. To fuse such complementarity, we propose a novel distillation-collaboration framework with two branches acting as CBP and VLP respectively. The framework is optimized through a dual-branch alternate training strategy. Specifically, during the B step, we distill the confident background pseudo-labels from the CBP branch; while during the F step, the confident foreground pseudo-labels are distilled from the VLP branch. As a result, the dualbranch complementarity is effectively fused to promote one strong alliance. Extensive experiments and ablation studies on THUMOS14 and ActivityNet1.2 reveal that our method significantly outperforms state-of-the-art methods.
Chen Ju, Kunhao Zheng, Jinxiang Liu, Peisen Zhao, Ya Zhang 0002, Jianlong Chang, Qi Tian 0001, Yanfeng Wang 0001
CVPR5
2023 Controllable Mesh Generation Through Sparse Latent Point Diffusion Models
abstract
Mesh generation is of great value in various applications involving computer graphics and virtual content, yet designing generative models for meshes is challenging due to their irregular data structure and inconsistent topology of meshes in the same category. In this work, we design a novel sparse latent point diffusion model for mesh generation. Our key insight is to regard point clouds as an intermediate representation of meshes, and model the distribution of point clouds instead. While meshes can be generated from point clouds via techniques like Shape as Points (SAP), the challenges of directly generating meshes can be effectively avoided. To boost the efficiency and controllability of our mesh generation method, we propose to further encode point clouds to a set of sparse latent points with pointwise semantic meaningful features, where two DDPMs are trained in the space of sparse latent points to respectively model the distribution of the latent point positions and features at these latent points. We find that sampling in this latent space is faster than directly sampling dense point clouds. Moreover, the sparse latent points also enable us to explicitly control both the overall structures and local details of the generated meshes. Extensive experiments are conducted on the ShapeNet dataset, where our proposed sparse latent point diffusion model achieves superior performance in terms of generation quality and controllability when compared to existing methods. Project page, code and appendix: https://slide-3d.github.io.
Zhaoyang Lyu, Jinyi Wang, Yuwei An, Ya Zhang 0002, Dahua Lin, Bo Dai 0002
CVPR4
2023 Class-Balancing Diffusion Models
abstract
Diffusion-based models have shown the merits of generating high-quality visual data while preserving better diversity in recent studies. However, such observation is only justified with curated data distribution, where the data samples are nicely pre-processed to be uniformly distributed in terms of their labels. In practice, a long-tailed data distribution appears more common and how diffusion models perform on such classimbalanced data remains unknown. In this work, we first investigate this problem and observe significant degradation in both diversity and fidelity when the diffusion model is trained on datasets with classimbalanced distributions. Especially in tail classes, the generations largely lose diversity and we observe severe mode-collapse issues. To tackle this problem, we set from the hypothesis that the data distribution is not class-balanced, and propose Class-Balancing Diffusion Models (CBDM) that are trained with a distribution adjustment regularizer as a solution. Experiments show that images generated by CBDM exhibit higher diversity and quality in both quantitative and qualitative ways. Our method benchmarked the generation results on CIFAR100/CIFAR100LT dataset and shows out-standing performance on the downstream recognition task.
Huangjie Zheng, Jiangchao Yao, Mingyuan Zhou, Ya Zhang 0002
CVPR5
2023 DR2: Diffusion-Based Robust Degradation Remover for Blind Face Restoration
abstract
Blind face restoration usually synthesizes degraded low-quality data with a pre-defined degradation model for training, while more complex cases could happen in the real world. This gap between the assumed and actual degradation hurts the restoration performance where artifacts are often observed in the output. However, it is expensive and infeasible to include every type of degradation to cover real-world cases in the training data. To tackle this robustness issue, we propose Diffusion-based Robust Degradation Remover (DR2) to first transform the degraded image to a coarse but degradation-invariant prediction, then employ an enhancement module to restore the coarse prediction to a high-quality image. By leveraging a well-performing denoising diffusion probabilistic model, our DR2 diffuses input images to a noisy status where various types of degradation give way to Gaussian noise, and then captures semantic information through iterative denoising steps. As a result, DR2 is robust against common degradation (e.g. blur, resize, noise and compression) and compatible with different designs of enhancement modules. Experiments in various settings show that our framework outperforms state-of-the-art methods on heavily degraded synthetic and real-world datasets.
Zhixin Wang, Ziying Zhang, Xiaoyun Zhang 0001, Huangjie Zheng, Mingyuan Zhou, Ya Zhang 0002, Yanfeng Wang 0001
CVPR6
2023 Federated Domain Generalization with Generalization Adjustment
abstract
Federated Domain Generalization (FedDG) attempts to learn a global model in a privacy-preserving manner that generalizes well to new clients possibly with domain shift. Recent exploration mainly focuses on designing an unbiased training strategy within each individual domain. However, without the support of multi-domain data jointly in the minibatch training, almost all methods cannot guarantee the generalization under domain shift. To overcome this problem, we propose a novel global objective incorporating a new variance reduction regularizer to encourage fairness. A novel FL-friendly method named Generalization Adjustment (GA) is proposed to optimize the above objective by dynamically calibrating the aggregation weights. The theoretical analysis of GA demonstrates the possibility to achieve a tighter generalization bound with an explicit reweighted aggregation, substituting the implicit multi-domain data sharing that is only applicable to the conventional DG settings. Besides, the proposed algorithm is generic and can be combined with any local client training-based methods. Extensive experiments on several benchmark datasets have shown the effectiveness of the proposed method, with consistent improvements over several FedDG algorithms when used in combination. The source code is released at https://github.com/MediaBrain-SJTU/FedDG-GA
Qinwei Xu, Jiangchao Yao, Ya Zhang 0002, Qi Tian 0001, Yanfeng Wang 0001
CVPR4
2023 Open-vocabulary Object Segmentation with Diffusion Models
abstract
The goal of this paper is to extract the visual-language correspondence from a pre-trained text-to-image diffusion model, in the form of segmentation map, i.e., simultaneously generating images and segmentation masks for the corresponding visual entities described in the text prompt. We make the following contributions: (i) we pair the existing Stable Diffusion model with a novel grounding module, that can be trained to align the visual and textual embedding space of the diffusion model with only a small number of object categories; (ii) we establish an automatic pipeline for constructing a dataset, that consists of {image, segmentation mask, text prompt} triplets, to train the proposed grounding module; (iii) we evaluate the performance of open-vocabulary grounding on images generated from the text-to-image diffusion model and show that the module can well segment the objects of categories beyond seen ones at training time, as shown in Fig. 1; (iv) we adopt the augmented diffusion model to build a synthetic semantic segmentation dataset, and show that, training a standard segmentation model on such dataset demonstrates competitive performance on the zero-shot segmentation (ZS3) benchmark, which opens up new opportunities for adopting the powerful diffusion model for discriminative tasks.
Ziyi Li 0004, Qinye Zhou, Xiaoyun Zhang 0001, Ya Zhang 0002, Yanfeng Wang 0001, Weidi Xie
ICCV4
2023 MedKLIP: Medical Knowledge Enhanced Language-Image Pre-Training for X-ray Diagnosis
abstract
In this paper, we consider enhancing medical visual-language pre-training (VLP) with domain-specific knowledge, by exploiting the paired image-text reports from the radiological daily practice. In particular, we make the following contributions: First, unlike existing works that directly process the raw reports, we adopt a novel triplet extraction module to extract the medical-related information, avoiding unnecessary complexity from language grammar and enhancing the supervision signals; Second, we propose a novel triplet encoding module with entity translation by querying a knowledge base, to exploit the rich domain knowledge in medical field, and implicitly build relationships between medical entities in the language embedding space; Third, we propose to use a Transformer-based fusion model for spatially aligning the entity description with visual signals at the image patch level, enabling the ability for medical diagnosis; Fourth, we conduct thorough experiments to validate the effectiveness of our architecture, and benchmark on numerous public benchmarks e.g., ChestX-ray14, RSNA Pneumonia, SIIM-ACR Pneumothorax, COVIDx CXR-2, COVID Rural, and EdemaSeverity. In both zero-shot and fine-tuning settings, our model has demonstrated strong performance compared with the former methods on disease classification and grounding.
Chaoyi Wu, Xiaoman Zhang, Ya Zhang 0002, Yanfeng Wang 0001, Weidi Xie
ICCV3
2023 Joint-Relation Transformer for Multi-Person Motion Prediction
abstract
Multi-person motion prediction is a challenging problem due to the dependency of motion on both individual past movements and interactions with other people. Transformer-based methods have shown promising results on this task, but they miss the explicit relation representation between joints, such as skeleton structure and pairwise distance, which is crucial for accurate interaction modeling. In this paper, we propose the Joint-Relation Transformer, which utilizes relation information to enhance interaction modeling and improve future motion prediction. Our relation information contains the relative distance and the intra-/inter-person physical constraints. To fuse relation and joint information, we design a novel joint-relation fusion layer with relation-aware attention to update both features. Additionally, we supervise the relation information by forecasting future distance. Experiments show that our method achieves a 13.4% improvement of 900ms VIM on 3DPW-SoMoF/RC and 17.8%/12.0% improvement of 3s MPJPE on CMU-Mpcap/MuPoTS-3D dataset. Code is available at https://github.com/MediaBrain-SJTU/JRTransformer.
Qingyao Xu, Weibo Mao, Jingze Gong, Chenxin Xu, Siheng Chen, Weidi Xie, Ya Zhang 0002, Yanfeng Wang 0001
ICCV7
2023 Long-Tailed Partial Label Learning via Dynamic Rebalancing
Feng Hong 0004, Jiangchao Yao, Zhihan Zhou 0002, Ya Zhang 0002, Yanfeng Wang 0001
ICLR4
2023 Multi-scale Cross-restoration Framework for Electrocardiogram Anomaly Detection
Aofan Jiang, Chaoqin Huang, Zi Zeng, Ya Zhang 0002, Yanfeng Wang 0001
MICCAI (1)7
2023 PMC-CLIP: Contrastive Language-Image Pre-training Using Biomedical Documents
Weixiong Lin, Xiaoman Zhang, Chaoyi Wu, Ya Zhang 0002, Yanfeng Wang 0001, Weidi Xie
MICCAI (8)5
2023 GRACE: A Generalized and Personalized Federated Learning Method for Medical Imaging
abstract
Federated learning has been extensively explored in privacy-preserving medical image analysis. However, the domain shift widely existed in real-world scenarios still greatly limits its practice, which requires to consider both generalization and personalization, namely generalized and personalized federated learning (GPFL). Previous studies almost focus on the partial objective of GPFL: personalized federated learning mainly cares about its local performance, which cannot guarantee a generalized global model for unseen clients; federated domain generalization only considers the out-of-domain performance, ignoring the performance of the training clients. To achieve both objectives effectively, we propose a novel GRAdient CorrEction (GRACE) method. GRACE incorporates a feature alignment regularization under a meta-learning framework on the client side to correct the personalized gradients from overfitting. Simultaneously, GRACE employs a consistency-enhanced re-weighting aggregation to calibrate the uploaded gradients on the server side for better generalization. Extensive experiments on two medical image benchmarks demonstrate the superiority of our method under various GPFL settings. Code available at https://github.com/MediaBrain-SJTU/GPFL-GRACE.
Ziqing Fan, Qinwei Xu, Jiangchao Yao, Ya Zhang 0002, Yanfeng Wang 0001
MICCAI (3)5
2023 Combating Representation Learning Disparity with Geometric Harmonization
abstract
Self-supervised learning (SSL) as an effective paradigm of representation learning has achieved tremendous success on various curated datasets in diverse scenarios. Nevertheless, when facing the long-tailed distribution in real-world applications, it is still hard for existing methods to capture transferable and robust representation. The attribution is that the vanilla SSL methods that pursue the sample-level uniformity easily leads to representation learning disparity, where head classes with the huge sample number dominate the feature regime but tail classes with the small sample number passively collapse. To address this problem, we propose a novel Geometric Harmonization (GH) method to encourage the category-level uniformity in representation learning, which is more benign to the minority and almost does not hurt the majority under long-tailed distribution. Specially, GH measures the population statistics of the embedding space on top of self-supervised learning, and then infer an fine-grained instance-wise calibration to constrain the space expansion of head classes and avoid the passive collapse of tail classes. Our proposal does not alter the setting of SSL and can be easily integrated into existing methods in a low-cost manner. Extensive results on a range of benchmark datasets show the effectiveness of \methodspace with high tolerance to the distribution skewness.
Zhihan Zhou 0002, Jiangchao Yao, Feng Hong 0004, Ya Zhang 0002, Bo Han 0003, Yanfeng Wang 0001
NeurIPS4
2023 Federated Learning with Bilateral Curation for Partially Class-Disjoint Data
abstract
Partially class-disjoint data (PCDD), a common yet under-explored data formation where each client contributes a part of classes (instead of all classes) of samples, severely challenges the performance of federated algorithms. Without full classes, the local objective will contradict the global objective, yielding the angle collapse problem for locally missing classes and the space waste problem for locally existing classes. As far as we know, none of the existing methods can intrinsically mitigate PCDD challenges to achieve holistic improvement in the bilateral views (both global view and local view) of federated learning. To address this dilemma, we are inspired by the strong generalization of simplex Equiangular Tight Frame (ETF) on the imbalanced data, and propose a novel approach called FedGELA where the classifier is globally fixed as a simplex ETF while locally adapted to the personal distributions. Globally, FedGELA provides fair and equal discrimination for all classes and avoids inaccurate updates of the classifier, while locally it utilizes the space of locally missing classes for locally existing classes. We conduct extensive experiments on a range of datasets to demonstrate that our FedGELA achieves promising performance (averaged improvement of 3.9% to FedAvg and 1.5% to best baselines) and provide both local and global convergence guarantees.
Ziqing Fan, Jiangchao Yao, Bo Han 0003, Ya Zhang 0002, Yanfeng Wang 0001
NeurIPS5
2023 Open-Vocabulary Semantic Segmentation via Attribute Decomposition-Aggregation
Chaofan Ma, Yuhuan Yang, Chen Ju, Fei Zhang 0016, Ya Zhang 0002, Yanfeng Wang 0001
NeurIPS5
2023 Asynchrony-Robust Collaborative Perception via Bird's Eye View Flow
abstract
Collaborative perception can substantially boost each agent's perception ability by facilitating communication among multiple agents. However, temporal asynchrony among agents is inevitable in the real world due to communication delays, interruptions, and clock misalignments. This issue causes information mismatch during multi-agent fusion, seriously shaking the foundation of collaboration. To address this issue, we propose CoBEVFlow, an asynchrony-robust collaborative perception system based on bird's eye view (BEV) flow. The key intuition of CoBEVFlow is to compensate motions to align asynchronous collaboration messages sent by multiple agents. To model the motion in a scene, we propose BEV flow, which is a collection of the motion vector corresponding to each spatial location. Based on BEV flow, asynchronous perceptual features can be reassigned to appropriate positions, mitigating the impact of asynchrony. CoBEVFlow has two advantages: (i) CoBEVFlow can handle asynchronous collaboration messages sent at irregular, continuous time stamps without discretization; and (ii) with BEV flow, CoBEVFlow only transports the original perceptual features, instead of generating new perceptual features, avoiding additional noises. To validate CoBEVFlow's efficacy, we create IRregular V2V(IRV2V), the first synthetic collaborative perception dataset with various temporal asynchronies that simulate different real-world scenarios. Extensive experiments conducted on both IRV2V and the real-world dataset DAIR-V2X show that CoBEVFlow consistently outperforms other baselines and is robust in extremely asynchronous settings. The code is available at https://github.com/MediaBrain-SJTU/CoBEVFlow.
Sizhe Wei, Yuxi Wei, Yue Hu 0011, Yiqi Zhong, Siheng Chen, Ya Zhang 0002
NeurIPS7
2023 Uncovering Prototypical Knowledge for Weakly Open-Vocabulary Semantic Segmentation
abstract
This paper studies the problem of weakly open-vocabulary semantic segmentation (WOVSS), which learns to segment objects of arbitrary classes using mere image-text pairs. Existing works turn to enhance the vanilla vision transformer by introducing explicit grouping recognition, i.e., employing several group tokens/centroids to cluster the image tokens and perform the group-text alignment. Nevertheless, these methods suffer from a granularity inconsistency regarding the usage of group tokens, which are aligned in the all-to-one v.s. one-to-one manners during the training and inference phases, respectively. We argue that this discrepancy arises from the lack of elaborate supervision for each group token. To bridge this granularity gap, this paper explores explicit supervision for the group tokens from the prototypical knowledge. To this end, this paper proposes the non-learnable prototypical regularization (NPR) where non-learnable prototypes are estimated from source features to serve as supervision and enable contrastive matching of the group tokens. This regularization encourages the group tokens to segment objects with less redundancy and capture more comprehensive semantic regions, leading to increased compactness and richness. Based on NPR, we propose the prototypical guidance segmentation network (PGSeg) that incorporates multi-modal regularization by leveraging prototypical sources from both images and texts at different levels, progressively enhancing the segmentation capability with diverse prototypical patterns. Experimental results show that our proposed method achieves state-of-the-art performance on several benchmark datasets.
Fei Zhang 0016, Tianfei Zhou, Boyang Li 0007, Chaofan Ma, Jiangchao Yao, Ya Zhang 0002, Yanfeng Wang 0001
NeurIPS8
2023 Collaborative Uncertainty Benefits Multi-Agent Multi-Modal Trajectory Forecasting
abstract
In multi-modal multi-agent trajectory forecasting, two major challenges have not been fully tackled: 1) how to measure the uncertainty brought by the interaction module that causes correlations among the predicted trajectories of multiple agents; 2) how to rank the multiple predictions and select the optimal predicted trajectory. In order to handle the aforementioned challenges, this work first proposes a novel concept, collaborative uncertainty (CU), which models the uncertainty resulting from interaction modules. Then we build a general CU-aware regression framework with an original permutation-equivariant uncertainty estimator to do both tasks of regression and uncertainty estimation. Furthermore, we apply the proposed framework to current SOTA multi-agent multi-modal forecasting systems as a plugin module, which enables the SOTA systems to: 1) estimate the uncertainty in the multi-agent multi-modal trajectory forecasting task; 2) rank the multiple predictions and select the optimal one based on the estimated uncertainty. We conduct extensive experiments on a synthetic dataset and two public large-scale multi-agent trajectory forecasting benchmarks. Experiments show that: 1) on the synthetic dataset, the CU-aware regression framework allows the model to appropriately approximate the ground-truth Laplace distribution; 2) on the multi-agent trajectory forecasting benchmarks, the CU-aware regression framework steadily helps SOTA systems improve their performances. Especially, the proposed framework helps VectorNet improve by 262 cm regarding the Final Displacement Error of the chosen optimal prediction on the nuScenes dataset; 3) in multi-agent multi-modal trajectory forecasting, prediction uncertainty is proportional to future stochasticity; 4) the estimated CU values are highly related to the interactive information among agents. The proposed framework can guide the development of more reliable and safer forecasting systems in the future.
Bohan Tang, Yiqi Zhong, Chenxin Xu, Ulrich Neumann, Ya Zhang 0002, Siheng Chen, Yanfeng Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 Latent Class-Conditional Noise Model
abstract
Learning with noisy labels has become imperative in the Big Data era, which saves expensive human labors on accurate annotations. Previous noise-transition-based methods have achieved theoretically-grounded performance under the Class-Conditional Noise model (CCN). However, these approaches builds upon an ideal but impractical anchor set available to pre-estimate the noise transition. Even though subsequent works adapt the estimation as a neural layer, the ill-posed stochastic learning of its parameters in back-propagation easily falls into undesired local minimums. We solve this problem by introducing a Latent Class-Conditional Noise model (LCCN) to parameterize the noise transition under a Bayesian framework. By projecting the noise transition into the Dirichlet space, the learning is constrained on a simplex characterized by the complete dataset, instead of some ad-hoc parametric space wrapped by the neural layer. We then deduce a dynamic label regression method for LCCN, whose Gibbs sampler allows us efficiently infer the latent true labels to train the classifier and to model the noise. Our approach safeguards the stable update of the noise transition, which avoids previous arbitrarily tuning from a mini-batch of samples. We further generalize LCCN to different counterparts compatible with open-set noisy labels, semi-supervised learning as well as cross-model training. A range of experiments demonstrate the advantages of LCCN and its variants over the current state-of-the-art methods. The code is available at here.
Jiangchao Yao, Bo Han 0003, Zhihan Zhou 0002, Ya Zhang 0002, Ivor W. Tsang
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Fourier-based augmentation with applications to domain generalization
Qinwei Xu, Ziqing Fan, Yanfeng Wang 0001, Yiyan Wu 0001, Ya Zhang 0002
Pattern Recognit.6
2023 Self-Supervised Tumor Segmentation With Sim2Real Adaptation
abstract
This paper targets on self-supervised tumor segmentation. We make the following contributions: (i) we take inspiration from the observation that tumors are often characterised independently of their contexts, we propose a novel proxy task "layer-decomposition", that closely matches the goal of the downstream task, and design a scalable pipeline for generating synthetic tumor data for pre-training; (ii) we propose a two-stage Sim2Real training regime for unsupervised tumor segmentation, where we first pre-train a model with simulated tumors, and then adopt a self-training strategy for downstream data adaptation; (iii) when evaluating on different tumor segmentation benchmarks, e.g. BraTS2018 for brain tumor segmentation and LiTS2017 for liver tumor segmentation, our approach achieves state-of-the-art segmentation performance under the unsupervised setting. While transferring the model for tumor segmentation under a low-annotation regime, the proposed approach also outperforms all existing self-supervised approaches; (iv) we conduct extensive ablation studies to analyse the critical components in data simulation, and validate the necessity of different proxy tasks. We demonstrate that, with sufficient texture randomization in simulation, model trained on synthetic data can effortlessly generalise to datasets with real tumors.
Xiaoman Zhang, Weidi Xie, Chaoqin Huang, Ya Zhang 0002, Xin Chen 0033, Qi Tian 0001, Yanfeng Wang 0001
IEEE J. Biomed. Health Informatics4
2023 Self-Supervised Masking for Unsupervised Anomaly Detection and Localization
abstract
Recently, anomaly detection and localization in multimedia data have received significant attention among the machine learning community. In real-world applications such as medical diagnosis and industrial defect detection, anomalies only present in a fraction of the images. To extend the reconstruction-based anomaly detection architecture to the localized anomalies, we propose a self-supervised learning approach throughrandom maskingand thenrestoring, namedSelf-SupervisedMasking(SSM) for unsupervised anomaly detection and localization. SSM not only enhances the training of the inpainting network but also leads to great improvement in the efficiency of mask prediction at inference. Through random masking, each image is augmented into a diverse set of training triplets, thus enabling the autoencoder to learn to reconstruct with masks of various sizes and shapes during training. To improve the efficiency and effectiveness of anomaly detection and localization at inference, we propose a novel progressive mask refinement approach that progressively uncovers the normal regions and finally locates the anomalous regions. The proposed SSM method outperforms several state-of-the-arts for both anomaly detection and anomaly localization, achieving 98.3% AUC on Retinal-OCT and 93.9% AUC on MVTec AD, respectively.
Chaoqin Huang, Qinwei Xu, Yanfeng Wang 0001, Yu Wang 0027, Ya Zhang 0002
IEEE Trans. Multim.5
2023 Adaptive Mutual Supervision for Weakly-Supervised Temporal Action Localization
abstract
Weakly-supervised temporal action localization aims to localize actions from untrimmed long videos with only video-level category labels. Most previous methods ignore the incompleteness issue of Class Activation Sequences (CAS), suffering from trivial detection results. To tackle this issue, we propose a novel Adaptive Mutual Supervision (AMS) framework with two branches, where the base branch detects the most discriminative action regions, while the supplementary branch localizes the less discriminative action regions through an adaptive sampler. The sampler dynamically updates the inputs for the supplementary branch using a sampling weight sequence negatively correlated with the CAS from the base branch, thus encouraging the supplementary branch to localize the action regions underestimated by the base branch. To promote mutual enhancement between two branches, we further construct mutual location supervision. Each branch adopts the location pseudo-labels generated from the other branch as the localization supervision. By alternately optimizing two branches for multiple iterations, we progressively complete action regions. Extensive experiments on THUMOS14 and ActivityNet1.2 demonstrate that the proposed AMS method significantly outperforms state-of-the-art methods.
Chen Ju, Peisen Zhao, Siheng Chen, Ya Zhang 0002, Xiaoyun Zhang 0001, Yanfeng Wang 0001, Qi Tian 0001
IEEE Trans. Multim.4
2023 Toward Equivalent Transformation of User Preferences in Cross Domain Recommendation
abstract
Cross domain recommendation (CDR) is one popular research topic in recommender systems. This article focuses on a popular scenario for CDR where different domains share the same set of users but no overlapping items. The majority of recent methods have explored the shared-user representation to transfer knowledge across domains. However, the idea of shared-user representation resorts to learning the overlapped features of user preferences and suppresses the domain-specific features. Other works try to capture the domain-specific features by an MLP mapping but require heuristic human knowledge of choosing samples to train the mapping. In this article, we attempt to learn both features of user preferences in a more principled way. We assume that each user’s preferences in one domain can be expressed by the other one, and these preferences can be mutually converted to each other with the so-called equivalent transformation. Based on this assumption, we propose an equivalent transformation learner (ETL), which models the joint distribution of user behaviors across domains. The equivalent transformation in ETL relaxes the idea of shared-user representation and allows the learned preferences in different domains to preserve the domain-specific features as well as the overlapped features. Extensive experiments on three public benchmarks demonstrate the effectiveness of ETL compared with recent state-of-the-art methods. Codes and data are available online: https://github.com/xuChenSJTU/ETL-master.
Xu Chen 0026, Ya Zhang 0002, Ivor W. Tsang, Yuangang Pan, Jingchao Su
ACM Trans. Inf. Syst.2
2022 Open-vocabulary Semantic Segmentation with Frozen Vision-Language Models
Chaofan Ma, Yuhuan Yang, Yanfeng Wang 0001, Ya Zhang 0002, Weidi Xie
BMVC4
2022 K-Space Transformer for Undersampled MRI Reconstruction
Weidi Xie, Yanfeng Wang 0001, Ya Zhang 0002
BMVC5
2022 A Simple Plugin for Transforming Images to Arbitrary Scales
Qinye Zhou, Ziyi Li 0004, Weidi Xie, Xiaoyun Zhang 0001, Yanfeng Wang 0001, Ya Zhang 0002
BMVC6
2022 LAR-SR: A Local Autoregressive Model for Image Super-Resolution
abstract
Previous super-resolution (SR) approaches often formulate SR as a regression problem and pixel wise restoration, which leads to a blurry and unreal SR output. Recent works combine adversarial loss with pixel-wise loss to train a GAN-based model or introduce normalizing flows into SR problems to generate more realistic images. As another powerful generative approach, autoregressive (AR) model has not been noticed in low level tasks due to its limitation. Based on the fact that given the structural in-formation, the textural details in the natural images are locally related without long term dependency, in this paper we propose a novel autoregressive model-based SR approach, namely LAR-SR, which can efficiently generate realistic SR images using a novel local autoregressive (LAR) module. The proposed LAR module can sample all the patches of textural components in parallel, which greatly reduces the time consumption. In addition to high time efficiency, it is also able to leverage contextual information of pixels and can be optimized with a consistent loss. Experimental results on the widely-used datasets show that the proposed LAR-SR approach achieves superior performance on the vi-sual quality and quantitative metrics compared with other generative models such as GAN, Flow, and is competitive with the mixture generative model.
Baisong Guo, Xiaoyun Zhang 0001, Haoning Wu 0002, Yu Wang 0027, Ya Zhang 0002, Yanfeng Wang 0001
CVPR5
2022 Task Decoupled Framework for Reference-based Super-Resolution
abstract
Reference-based super-resolution(RefSR) has achieved impressive progress on the recovery of high-frequency details thanks to an additional reference high-resolution(HR) image input. Although the superiority compared with Single-Image Super-Resolution(SISR), existing RefSR methods easily result in the reference-underuse issue and the reference-misuse as shown in Fig. I. In this work, we deeply investigate the cause of the two issues and further propose a novel framework to mitigate them. Our studies find that the issues are mostly due to the improper coupled framework design of current methods. Those methods conduct the super-resolution task of the input low-resolution(LR) image and the texture transfer task from the reference image together in one module, easily introducing the interference between LR and reference features. Inspired by this finding, we propose a novel framework, which decouples the two tasks of RefSR, eliminating the interference between the LR image and the reference image. The super-resolution task upsamples the LR image leveraging only the LR image itself. The texture transfer task extracts and transfers abundant textures from the reference image to the coarsely upsampled result of the super-resolution task. Extensive experiments demonstrate clear improvements in both quantitative and qualitative evaluations over state-of-the-art methods.
Xiaoyun Zhang 0001, Siheng Chen, Ya Zhang 0002, Yanfeng Wang 0001, Dazhi He
CVPR5
2022 GroupNet: Multiscale Hypergraph Neural Networks for Trajectory Prediction with Relational Reasoning
abstract
Demystifying the interactions among multiple agents from their past trajectories is fundamental to precise and interpretable trajectory prediction. However, previous works only consider pair-wise interactions with limited relational reasoning. To promote more comprehensive interaction modeling for relational reasoning, we propose GroupNet, a multiscale hypergraph neural network, which is novel in terms of both interaction capturing and representation learning. From the aspect of interaction capturing, we propose a trainable multiscale hypergraph to capture both pair-wise and group-wise interactions at multiple group sizes. From the aspect of interaction representation learning, we propose a three-element format that can be learnt end-to-end and explicitly reason some relational factors including the interaction strength and category. We apply GroupNet into both CVAE-based prediction system and previous state-of-the-art prediction systems for predicting socially plausible trajectories with relational reasoning. To validate the ability of relational reasoning, we experiment with synthetic physics simulations to reflect the ability to capture group behaviors, reason interaction strength and interaction category. To validate the effectiveness of prediction, we conduct extensive experiments on three real-world trajectory prediction datasets, including NBA, SDD and ETH-UCY; and we show that with GroupNet, the CVAE-based prediction system outperforms state-of-the-art methods. We also show that adding GroupNet will further improve the performance of previous state-of-the-art prediction systems.
Chenxin Xu, Maosen Li, Zhenyang Ni, Ya Zhang 0002, Siheng Chen
CVPR4
2022 Registration Based Few-Shot Anomaly Detection
Chaoqin Huang, Haoyan Guan, Aofan Jiang, Ya Zhang 0002, Michael W. Spratling, Yanfeng Wang 0001
ECCV (24)4
2022 Prompting Visual-Language Models for Efficient Video Understanding
Chen Ju, Tengda Han, Kunhao Zheng, Ya Zhang 0002, Weidi Xie
ECCV (35)4
2022 Skeleton-Parted Graph Scattering Networks for 3D Human Motion Prediction
Maosen Li, Siheng Chen, Lingxi Xie, Qi Tian 0001, Ya Zhang 0002
ECCV (6)6
2022 Spatio-Temporal Graph Complementary Scattering Networks
abstract
Spatio-temporal graph signal analysis has a significant impact on a wide range of applications, including hand/body pose action recognition. To achieve effective analysis, spatio-temporal graph convolutional networks (ST-GCN) leverage the powerful learning ability to achieve great empirical successes; however, those methods need a huge amount of high-quality training data and lack theoretical interpretation. To address this issue, the spatio-temporal graph scattering transform (ST-GST) was proposed to put forth a theoretically interpretable framework; however, the empirical performance of this approach is constrained by the fully mathematical design. To benefit from both sides, this work proposes a novel complementary mechanism to organically combine the spatio-temporal graph scattering transform and neural networks, resulting in the proposed spatio-temporal graph complementary scattering networks (ST-GCSN). The essence is to leverage the mathematically designed graph wavelets with pruning techniques to cover major information and use trainable networks to capture complementary information. The empirical experiments on hand pose action recognition show that the proposed ST-GCSN outperforms both ST-GCN and ST-GST.
Zida Cheng, Siheng Chen, Ya Zhang 0002
ICASSP3
2022 FedSkip: Combatting Statistical Heterogeneity with Federated Skip Aggregation
abstract
The statistical heterogeneity of the non-independent and identically distributed (non-IID) data in local clients significantly limits the performance of federated learning. Previous attempts like FedProx, SCAFFOLD, MOON, FedNova and FedDyn resort to an optimization perspective, which requires an auxiliary term or re-weights local updates to calibrate the learning bias or the objective inconsistency. However, in addition to previous explorations for improvement in federated averaging, our analysis shows that another critical bottleneck is the poorer optima of client models in more heterogeneous conditions. We thus introduce a data-driven approach called FedSkip to improve the client optima by periodically skipping federated averaging and scattering local models to the cross devices. We provide theoretical analysis of the possible benefit from FedSkip and conduct extensive experiments on a range of datasets to demonstrate that FedSkip achieves much higher accuracy, better aggregation efficiency and competing communication efficiencys. Source code is available at: https://github.com/MediaBrain-SJTU/FedSkip.
Ziqing Fan, Yanfeng Wang 0001, Jiangchao Yao, Lingjuan Lyu, Ya Zhang 0002, Qi Tian 0001
ICDM5
2022 Contrastive Learning with Boosted Memorization
abstract
Self-supervised learning has achieved a great success in the representation learning of visual and textual data. However, the current methods are mainly validated on the well-curated datasets, which do not exhibit the real-world long-tailed distribution. Recent attempts to consider self-supervised long-tailed learning are made by rebalancing in the loss perspective or the model perspective, resembling the paradigms in the supervised long-tailed learning. Nevertheless, without the aid of labels, these explorations have not shown the expected significant promise due to the limitation in tail sample discovery or the heuristic structure design. Different from previous works, we explore this direction from an alternative perspective, i.e., the data perspective, and propose a novel Boosted Contrastive Learning (BCL) method. Specifically, BCL leverages the memorization effect of deep neural networks to automatically drive the information discrepancy of the sample views in contrastive learning, which is more efficient to enhance the long-tailed learning in the label-unaware context. Extensive experiments on a range of benchmark datasets demonstrate the effectiveness of BCL over several state-of-the-art methods. Our code is available at https://github.com/MediaBrain-SJTU/BCL.
Zhihan Zhou 0002, Jiangchao Yao, Yanfeng Wang 0001, Bo Han 0003, Ya Zhang 0002
ICML5
2022 Boundary-Enhanced Self-supervised Learning for Brain Structure Segmentation
Feng Chang, Chaoyi Wu, Yanfeng Wang 0001, Ya Zhang 0002, Xin Chen 0033, Qi Tian 0001
MICCAI (1)4
2022 Transforming the Interactive Segmentation for Medical Imaging
Chaofan Ma, Yuhuan Yang, Weidi Xie, Ya Zhang 0002
MICCAI (4)5
2022 Exploiting Transformation Invariance and Equivariance for Self-supervised Sound Localisation
abstract
We present a simple yet effective self-supervised framework for audio-visual representation learning, to localize the sound source in videos. To understand what enables to learn useful representations, we systematically investigate the effects of data augmentations, and reveal that (1) composition of data augmentations plays a critical role, i.e. explicitly encouraging the audio-visual representations to be invariant to various transformations (transformation invariance); (2) enforcing geometric consistency substantially improves the quality of learned representations, i.e. the detected sound source should follow the same transformation applied on input video frames (transformation equivariance). Extensive experiments demonstrate that our model significantly outperforms previous methods on two sound localization benchmarks, namely, Flickr-SoundNet and VGG-Sound. Additionally, we also evaluate audio retrieval and cross-modal retrieval tasks. In both cases, our self-supervised models demonstrate superior retrieval performances, even competitive with the supervised approach in audio retrieval. This reveals the proposed framework learns strong multi-modal representations that are beneficial to sound localisation and generalization to further applications. The project page is https://jinxiang-liu.github.io/SSL-TIE.
Jinxiang Liu, Chen Ju, Weidi Xie, Ya Zhang 0002
ACM Multimedia4
2022 TED: Two-stage expert-guided interpretable diagnosis framework for microvascular invasion in hepatocellular carcinoma
Shuwen Sun, Qiu-Ping Liu, Ya Zhang 0002
Medical Image Anal.5
2022 Learning on Attribute-Missing Graphs
abstract
Graphs with complete node attributes have been widely explored recently. While in practice, there is a graph where attributes of only partial nodes could be available and those of the others might be entirely missing. This attribute-missing graph is related to numerous real-world applications and there are limited studies investigating the corresponding learning problems. Existing graph learning methods including the popular GNN cannot provide satisfied learning performance since they are not specified for attribute-missing graphs. Thereby, designing a new GNN for these graphs is a burning issue to the graph learning community. In this article, we make a shared-latent space assumption on graphs and develop a novel distribution matching-based GNN called structure-attribute transformer (SAT) for attribute-missing graphs. SAT leverages structures and attributes in a decoupled scheme and achieves the joint distribution modeling of structures and attributes by distribution matching techniques. It could not only perform the link prediction task but also the newly introduced node attribute completion task. Furthermore, practical measures are introduced to quantify the performance of node attribute completion. Extensive experiments on seven real-world datasets indicate SAT shows better performance than other methods on both link prediction and node attribute completion tasks.
Xu Chen 0026, Siheng Chen, Jiangchao Yao, Huangjie Zheng, Ya Zhang 0002, Ivor W. Tsang
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Symbiotic Graph Neural Networks for 3D Skeleton-Based Human Action Recognition and Motion Prediction
abstract
3D skeleton-based action recognition and motion prediction are two essential problems of human activity understanding. In many previous works: 1) they studied two tasks separately, neglecting internal correlations; and 2) they did not capture sufficient relations inside the body. To address these issues, we propose a symbiotic model to handle two tasks jointly; and we propose two scales of graphs to explicitly capture relations among body-joints and body-parts. Together, we propose symbiotic graph neural networks, which contain a backbone, an action-recognition head, and a motion-prediction head. Two heads are trained jointly and enhance each other. For the backbone, we propose multi-branch multiscale graph convolution networks to extract spatial and temporal features. The multiscale graph convolution networks are based on joint-scale and part-scale graphs. The joint-scale graphs contain actional graphs, capturing action-based relations, and structural graphs, capturing physical constraints. The part-scale graphs integrate body-joints to form specific parts, representing high-level relations. Moreover, dual bone-based graphs and networks are proposed to learn complementary features. We conduct extensive experiments for skeleton-based action recognition and motion prediction with four datasets, NTU-RGB+D, Kinetics, Human3.6M, and CMU Mocap. Experiments show that our symbiotic graph neural networks achieve better performances on both tasks compared to the state-of-the-art methods.
Maosen Li, Siheng Chen, Xu Chen 0026, Ya Zhang 0002, Yanfeng Wang 0001, Qi Tian 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Progressive privileged knowledge distillation for online action detection
Peisen Zhao, Lingxi Xie, Ya Zhang 0002, Qi Tian 0001
Pattern Recognit.4
2022 Actionness-Guided Transformer for Anchor-Free Temporal Action Localization
abstract
Temporal action localization, detecting actions in untrimmed videos, is widely studied by anchor-based approaches that first generate excessive action proposals,i.e., temporal windows, then evaluate and classify these proposals. To reduce the number of action proposals, recent studies use an anchor-free approach that leverages each time point rather than a temporal window to represent an action instance. However, this point representation, usually modeled by temporal convolutions, may have the fixed and limited receptive field to detect an entire action. So we propose an Actionness-guided Transformer (Ag-Trans) model to learn representations for each point proposal. Ag-Trans first predicts the actionness,i.e., time sequences of the action starting, continuing, and ending phases, then the corresponding action phase can be embedded to model the point representation. Experimental results show that the Ag-Trans model outperforms the CNN-based model under the same experiment settings, especially for long-duration actions.
Peisen Zhao, Lingxi Xie, Ya Zhang 0002, Qi Tian 0001
IEEE Signal Process. Lett.3
2022 A 3D Mesh-Based Lifting-and-Projection Network for Human Pose Transfer
abstract
Human pose transfer has typically been modeled as a 2D image-to-image translation problem. This formulation ignores the human body shape prior in 3D space and inevitably causes implausible artifacts, especially when facing occlusion. To address this issue, we propose alifting-and-projectionframework to perform pose transfer in the 3D mesh space. The core of our framework is a foreground generation module, that consists of two novel networks: a lifting-and-projection network (LPNet) and an appearance detail compensating network (ADCNet). To leverage the human body shape prior, LPNet exploits the topological information of the body mesh to learn an expressive visual representation for the target person in the 3D mesh space. To preserve texture details, ADCNet is further introduced to enhance the feature produced by LPNet with the source foreground image. Such design of the foreground generation module enables the model to better handle difficult cases such as those with occlusions. Experiments on the iPER and Fashion datasets empirically demonstrate that the proposed lifting-and-projection framework is effective and outperforms the existing image-to-image-based and mesh-based methods on human pose transfer task in both self-transfer and cross-transfer settings.
Jinxiang Liu, Yangheng Zhao, Siheng Chen, Ya Zhang 0002
IEEE Trans. Multim.4
2022 Attribute Restoration Framework for Anomaly Detection
abstract
With the recent advances in deep neural networks, anomaly detection in multimedia has received much attention in the computer vision community. While reconstruction-based methods have recently shown great promise for anomaly detection, the information equivalence among input and supervision for reconstruction tasks can not effectively force the network to learn semantic feature embeddings. We here propose to break this equivalence by erasing selected attributes from the original data and reformulate it as a restoration task, where the normal and the anomalous data are expected to be distinguishable based on restoration errors. Through forcing the network to restore the original image, the semantic feature embeddings related to the erased attributes are learned by the network. During testing phases, because anomalous data are restored with the attribute learned from the normal data, the restoration error is expected to be large. Extensive experiments have demonstrated that the proposed method significantly outperforms several state-of-the-arts on multiple benchmark datasets, especially on ImageNet, increasing the AUROC of the top-performing baseline by 10.1%. We also evaluate our method on a real-world anomaly detection dataset MVTec AD.
Chaoqin Huang, Jinkun Cao, Maosen Li, Ya Zhang 0002, Cewu Lu
IEEE Trans. Multim.5
2021 Invariant Teacher and Equivariant Student for Unsupervised 3D Human Pose Estimation
abstract
We propose a novel method based on teacher-student learning framework for 3D human pose estimation without any 3D annotation or side information. To solve this unsupervised-learning problem, the teacher network adopts pose-dictionary-based modeling for regularization to estimate a physically plausible 3D pose. To handle the decomposition ambiguity in the teacher network, we propose a cycle-consistent architecture promoting a 3D rotation-invariant property to train the teacher network. To further improve the estimation accuracy, the student network adopts a novel graph convolution network for flexibility to directly estimate the 3D coordinates. Another cycle-consistent architecture promoting 3D rotation-equivariant property is adopted to exploit geometry consistency, together with knowledge distillation from the teacher network to improve the pose estimation performance. We conduct extensive experiments on Human3.6M and MPI-INF-3DHP. Our method reduces the 3D joint prediction error by 11.4% compared to state-of-the-art unsupervised methods and also outperforms many weakly-supervised methods that use side information on Human3.6M. Code will be available at https://github.com/sjtuxcx/ITES.
Chenxin Xu, Siheng Chen, Maosen Li, Ya Zhang 0002
AAAI4
2021 ESAD: End-to-end Semi-supervised Anomaly Detection
Chaoqin Huang, Peisen Zhao, Ya Zhang 0002, Yanfeng Wang 0001, Qi Tian 0001
BMVC4
2021 A Fourier-Based Framework for Domain Generalization
abstract
Modern deep neural networks suffer from performance degradation when evaluated on testing data under different distributions from training data. Domain generalization aims at tackling this problem by learning transferable knowledge from multiple source domains in order to generalize to unseen target domains. This paper introduces a novel Fourier-based perspective for domain generalization. The main assumption is that the Fourier phase information contains high-level semantics and is not easily affected by domain shifts. To force the model to capture phase information, we develop a novel Fourier-based data augmentation strategy called amplitude mix which linearly interpolates between the amplitude spectrums of two images. A dual-formed consistency loss called co-teacher regularization is further introduced between the predictions induced from original and augmented images. Extensive experiments on three benchmarks have demonstrated that the proposed method is able to achieve state-of-the-arts performance for domain generalization.
Qinwei Xu, Ya Zhang 0002, Yanfeng Wang 0001, Qi Tian 0001
CVPR3
2021 CaT: Weakly Supervised Object Detection with Category Transfer
abstract
A large gap exists between fully-supervised object detection and weakly-supervised object detection. To narrow this gap, some methods consider knowledge transfer from additional fully-supervised dataset. But these methods do not fully exploit discriminative category information in the fully-supervised dataset, thus causing low mAP. To solve this issue, we propose a novel category transfer framework for weakly supervised object detection. The intuition is to fully leverage both visually-discriminative and semantically-correlated category information in the fully-supervised dataset to enhance the object-classification ability of a weakly-supervised detector. To handle overlapping category transfer, we propose a double-supervision mean teacher to gather common category information and bridge the domain gap between two datasets. To handle non-overlapping category transfer, we propose a semantic graph convolutional network to promote the aggregation of semantic features between correlated categories. Experiments are conducted with Pascal VOC 2007 as the target weakly-supervised dataset and COCO as the source fully-supervised dataset. Our category transfer framework achieves 63.5% mAP and 80.3% CorLoc with 5 overlapping categories between two datasets, which outperforms the state-of-the-art methods. Codes are avaliable at https://github.com/MediaBrain-SJTU/CaT.
Lianyu Du, Xiaoyun Zhang 0001, Siheng Chen, Ya Zhang 0002, Yanfeng Wang 0001
ICCV5
2021 Divide and Conquer for Single-frame Temporal Action Localization
abstract
Single-frame temporal action localization (STAL) aims to localize actions in untrimmed videos with only one timestamp annotation for each action instance. Existing methods adopt the one-stage framework but couple the counting goal and the localization goal. This paper proposes a novel two-stage framework for the STAL task with the spirit of divide and conquer. The instance counting stage leverages the location supervision to determine the number of action instances and divide a whole video into multiple video clips, so that each video clip contains only one complete action instance; and the location estimation stage leverages the category supervision to localize the action instance in each video clip. To efficiently represent the action instance in each video clip, we introduce the proposal-based representation, and design a novel differentiable mask generator to enable the end-to-end training supervised by category labels. On THUMOS14, GTEA, and BEOID datasets, our method outperforms state-of-the-art methods by 3.5%, 2.7%, 4.8% mAP on average. And extensive experiments verify the effectiveness of our method.
Chen Ju, Peisen Zhao, Siheng Chen, Ya Zhang 0002, Yanfeng Wang 0001, Qi Tian 0001
ICCV4
2021 Semi-Supervised 3D Hand-Object Pose Estimation Via Pose Dictionary Learning
abstract
3D hand-object pose estimation is an important issue to understand the interaction between human and environment. Current hand-object pose estimation methods require detailed 3D labels, which are expensive and labor-intensive. To tackle the problem of data collection, we propose a semi-supervised 3D hand-object pose estimation method with two key techniques: pose dictionary learning and an object-oriented coordinate system. The proposed pose dictionary learning module can distinguish infeasible poses by reconstruction error, enabling unlabeled data to provide supervision signals. The proposed object-oriented coordinate system can make 3D estimations equivariant to the camera perspective. Experiments are conducted on FPHA and HO-3D datasets. Our method reduces estimation error by 19.5% / 24.9% for hands/objects compared to straightforward use of labeled data on FPHA and outperforms several baseline methods. Extensive experiments also validate the robustness of the proposed method.
Zida Cheng, Siheng Chen, Ya Zhang 0002
ICIP3
2021 Unsupervised Segmentation Framework with Active Contour Models for Cine Cardiac MRI
abstract
Deep learning methods have made remarkable progress in medical image segmentation tasks, but these methods require enough labeled data, which tends to be difficult for medical tasks. To tack this issue, we propose an unsupervised segmentation framework by combining deep learning networks with the active contour model. We design an iterative loop process that the network can be trained with outputs from the active contour model and the active contour model can be initialized by coarse predictions from the network. In this way, our approach can train the segmentation networks iterative with no annotations but only one initialization for the active contour model at the beginning. We evaluate our approach in the task of cine cardiac MRI segmentation and get very competitive results.
Lianyu Du, Xiaoyun Zhang 0001, Yu-Min Zhong, Ya Zhang 0002, Yanfeng Wang 0001
ICIP5
2021 Deep Unsupervised Image Anomaly Detection: An Information Theoretic Framework
abstract
Surrogate task based methods have recently shown great promise for unsupervised image anomaly detection. However, there is no guarantee that the surrogate tasks share the consistent optimization direction with anomaly detection. In this paper, we return to a direct objective function for anomaly detection with information theory, which maximizes the distance between normal and anomalous data in terms of the joint distribution of images and their representation. To make this objective function directly optimizable under the unsupervised setting, we manage to find its lower bound which weights the trade-off between mutual information and entropy, which leads to a novel information theoretic framework for unsupervised image anomaly detection. Extensive experiments on several benchmark data sets have shown that the proposed framework significantly outperforms several state-of-the-arts.
Huangjie Zheng, Chaoqin Huang, Ya Zhang 0002
ICIP4
2021 Cooperative Learning for Noisy Supervision
abstract
Learning with noisy labels has gained the enormous interest in the robust deep learning area. Recent studies have empir-ically disclosed that utilizing dual networks can enhance the performance of single network but without theoretic proof. In this paper, we propose Cooperative Learning (CooL) frame-work for noisy supervision that analytically explains the ef-fects of leveraging dual or multiple networks. Specifically, the simple but efficient combination in CooL yields a more reliable risk minimization for unseen clean data. A range of experiments have been conducted on several benchmarks with both synthetic and real-world settings. Extensive results indi-cate that CooL outperforms several state-of-the-art methods.
Hao Wu 0075, Jiangchao Yao, Ya Zhang 0002, Yanfeng Wang 0001
ICME3
2021 Sequential Learning on Liver Tumor Boundary Semantics and Prognostic Biomarker Mining
Jieneng Chen, Ke Yan 0006, Youbao Tang, Shuwen Sun, Qiuping Liu, Lingyun Huang, Jing Xiao 0006, Alan L. Yuille, Ya Zhang 0002, Le Lu 0001
MICCAI (7)11
2021 Fully Test-Time Adaptation for Image Segmentation
Minhao Hu, Tao Song 0002, Yujun Gu, Xiangde Luo, Jieneng Chen, Ya Zhang 0002, Shaoting Zhang 0001
MICCAI (3)7
2021 SAR: Scale-Aware Restoration Learning for 3D Tumor Segmentation
Xiaoman Zhang, Shixiang Feng, Ya Zhang 0002, Yanfeng Wang 0001
MICCAI (2)4
2021 Collaborative Uncertainty in Multi-Agent Trajectory Forecasting
abstract
Uncertainty modeling is critical in trajectory-forecasting systems for both interpretation and safety reasons. To better predict the future trajectories of multiple agents, recent works have introduced interaction modules to capture interactions among agents. This approach leads to correlations among the predicted trajectories. However, the uncertainty brought by such correlations is neglected. To fill this gap, we propose a novel concept, collaborative uncertainty (CU), which models the uncertainty resulting from the interaction module. We build a general CU-based framework to make a prediction model learn the future trajectory and the corresponding uncertainty. The CU-based framework is integrated as a plugin module to current state-of-the-art (SOTA) systems and deployed in two special cases based on multivariate Gaussian and Laplace distributions. In each case, we conduct extensive experiments on two synthetic datasets and two public, large-scale benchmarks of trajectory forecasting. The results are promising: 1) The results of synthetic datasets show that CU-based framework allows the model to nicely rebuild the ground-truth distribution. 2) The results of trajectory forecasting benchmarks demonstrate that the CU-based framework steadily helps SOTA systems improve their performances. Specially, the proposed CU-based framework helps VectorNet improve by 57 cm regarding Final Displacement Error on nuScenes dataset. 3) The visualization results of CU illustrate that the value of CU is highly related to the amount of the interactive information among agents.
Bohan Tang, Yiqi Zhong, Ulrich Neumann, Siheng Chen, Ya Zhang 0002
NeurIPS6
2021 Handwritten Chinese Font Generation with Collaborative Stroke Refinement
abstract
Automatic character generation is an appealing solution for typeface design, especially for Chinese fonts with over 3700 most commonly-used characters. This task is particularly challenging for handwritten characters with thin strokes which are error-prone during deformation. To handle the generation of thin strokes, we introduce an auxiliary branch for stroke refinement. The auxiliary branch is trained to generate the bold version of target characters which are then fed to the dominating branch to guide the stroke refinement. The two branches are jointly trained in a collaborative fashion. In addition, for practical use, it is desirable to train the character synthesis model with a small set of manually designed characters. Taking advantage of content-reuse phenomenon in Chinese characters, we further propose an online zoom-augmentation strategy to reduce the dependency on large size training sets. The proposed model is trained end-to-end and can be added on top of any method for font synthesis. Experimental results on handwritten font synthesis have shown that the proposed method significantly outperforms the state-of-the-art methods under practical setting, i.e. with only 750 paired training samples.
Chuan Wen, Yujie Pan 0001, Ya Zhang 0002, Siheng Chen, Yanfeng Wang 0001, Qi Tian 0001
WACV4
2021 Multiscale Spatio-Temporal Graph Neural Networks for 3D Skeleton-Based Motion Prediction
abstract
We propose a multiscale spatio-temporal graph neural network (MST-GNN) to predict the future 3D skeleton-based human poses in an action-category-agnostic manner. The core of MST-GNN is a multiscale spatio-temporal graph that explicitly models the relations in motions at various spatial and temporal scales. Different from many previous hierarchical structures, our multiscale spatio-temporal graph is built in a data-adaptive fashion, which captures nonphysical, yet motion-based relations. The key module of MST-GNN is a multiscale spatio-temporal graph computational unit (MST-GCU) based on the trainable graph structure. MST-GCU embeds underlying features at individual scales and then fuses features across scales to obtain a comprehensive representation. The overall architecture of MST-GNN follows an encoder-decoder framework, where the encoder consists of a sequence of MST-GCUs to learn the spatial and temporal features of motions, and the decoder uses a graph-based attention gate recurrent unit (GA-GRU) to generate future poses. Extensive experiments are conducted to show that the proposed MST-GNN outperforms state-of-the-art methods in both short and long-term motion prediction on the datasets of Human 3.6M, CMU Mocap and 3DPW, where MST-GNN outperforms previous works by 5.33% and 3.67% of mean angle errors in average for short-term and long-term prediction on Human 3.6M, and by 11.84% and 4.71% of mean angle errors for short-term and long-term prediction on CMU Mocap, and by 1.13% of mean angle errors on 3DPW in average, respectively. We further investigate the learned multiscale graphs for interpretability.
Maosen Li, Siheng Chen, Yangheng Zhao, Ya Zhang 0002, Yanfeng Wang 0001, Qi Tian 0001
IEEE Trans. Image Process.4
2021 Two-Stream Compare and Contrast Network for Vertebral Compression Fracture Diagnosis
abstract
Differentiating Vertebral Compression Fractures (VCFs) associated with trauma and osteoporosis (benign VCFs) or those caused by metastatic cancer (malignant VCFs) is critically important for treatment decisions. So far, automatic VCFs diagnosis is solved in a two-step manner, i.e., first identify VCFs and then classify them into benign or malignant. In this paper, we explore to model VCFs diagnosis as a three-class classification problem, i.e., normal vertebrae, benign VCFs, and malignant VCFs. However, VCFs recognition and classification require very different features, and both tasks are characterized by high intra-class variation and high inter-class similarity. Moreover, the dataset is extremely class-imbalanced. To address the above challenges, we propose a novel Two-Stream Compare and Contrast Network (TSCCN) for VCFs diagnosis. This network consists of two streams, a recognition stream which learns to identify VCFs through comparing and contrasting between adjacent vertebrae, and a classification stream which compares and contrasts between intra-class and inter-class to learn features for fine-grained classification. The two streams are integrated via a learnable weight control module which adaptively sets their contribution. TSCCN is evaluated on a dataset consisting of 239 VCFs patients and achieves the average sensitivity and specificity of 92.56% and 96.29%, respectively.
Shixiang Feng, Ya Zhang 0002, Xiaoyun Zhang 0001
IEEE Trans. Medical Imaging3
2021 Boundary-Aware Supervoxel-Level Iteratively Refined Interactive 3D Image Segmentation With Multi-Agent Reinforcement Learning
abstract
Interactive segmentation has recently been explored to effectively and efficiently harvest high-quality segmentation masks by iteratively incorporating user hints. While iterative in nature, most existing interactive segmentation methods tend to ignore the dynamics of successive interactions and take each interaction independently. We here propose to model iterative interactive image segmentation with a Markov decision process (MDP) and solve it with reinforcement learning (RL) where each voxel is treated as an agent. Considering the large exploration space for voxel-wise prediction and the dependence among neighboring voxels for the segmentation tasks, multi-agent reinforcement learning is adopted, where the voxel-level policy is shared among agents. Considering that boundary voxels are more important for segmentation, we further introduce a boundary-aware reward, which consists of a global reward in the form of relative cross-entropy gain, to update the policy in a constrained direction, and a boundary reward in the form of relative weight, to emphasize the correctness of boundary predictions. To combine the advantages of different types of interactions, i. e., simple and efficient for point-clicking, and stable and robust for scribbles, we propose a supervoxel-clicking based interaction design. Experimental results on four benchmark datasets have shown that the proposed method significantly outperforms the state-of-the-arts, with the advantage of fewer interactions, higher accuracy, and enhanced robustness.
Chaofan Ma, Qisen Xu, Xiangfeng Wang 0001, Bo Jin 0003, Xiaoyun Zhang 0001, Yanfeng Wang 0001, Ya Zhang 0002
IEEE Trans. Medical Imaging7
2021 Universal-to-Specific Framework for Complex Action Recognition
abstract
Video-based action recognition has recently attracted much attention in the field of computer vision. To solve more complex recognition tasks, it has become necessary to distinguish different levels of interclass variations. Inspired by a common flowchart based on the human decision-making process that first narrows down the probable classes and then applies a "rethinking" process for finer-level recognition, we propose an effective universal-to-specific (U2S) framework for complex action recognition. The U2S framework is composed of three subnetworks: a universal network, a category-specific network, and a mask network. The universal network first learns universal feature representations. The mask network then generates attention masks for confusing classes through category regularization based on the output of the universal network. The mask is further used to guide the category-specific network for class-specific feature representations. The entire framework is optimized in an end-to-end manner. Experiments on a variety of benchmark datasets, e.g., the Something-Something, UCF101, and HMDB51 datasets, demonstrate the effectiveness of the U2S framework; i.e., U2S can focus on discriminative spatiotemporal regions for confusing categories. We further visualize the relationship between different classes, showing that U2S indeed improves the discriminability of learned features. Moreover, the proposed U2S model is a general framework and may adopt any base recognition network.
Peisen Zhao, Lingxi Xie, Ya Zhang 0002, Qi Tian 0001
IEEE Trans. Multim.3
2021 Decoupled Variational Embedding for Signed Directed Networks
abstract
Node representation learning for signed directed networks has received considerable attention in many real-world applications such as link sign prediction, node classification, and node recommendation. The challenge lies in how to adequately encode the complex topological information of the networks. Recent studies mainly focus on preserving the first-order network topology that indicates the closeness relationships of nodes. However, these methods generally fail to capture the high-order topology that indicates the local structures of nodes and serves as an essential characteristic of the network topology. In addition, for the first-order topology, the additional value of non-existent links is largely ignored. In this article, we propose to learn more representative node embeddings by simultaneously capturing the first-order and high-order topology in signed directed networks. In particular, we reformulate the representation learning problem on signed directed networks from a variational auto-encoding perspective and further develop a decoupled variational embedding (DVE) method. DVE leverages a specially designed auto-encoder structure to capture both the first-order and high-order topology of signed directed networks, and thus learns more representative node embeddings. Extensive experiments are conducted on three widely used real-world datasets. Comprehensive results on both link sign prediction and node recommendation task demonstrate the effectiveness of DVE. Qualitative results and analysis are also given to provide a better understanding of DVE.
Xu Chen 0026, Jiangchao Yao, Maosen Li, Ya Zhang 0002, Yanfeng Wang 0001
ACM Trans. Web4
2020 From Quantized DNNs to Quantizable DNNs
Kunyuan Du, Ya Zhang 0002, Haibing Guan
BMVC2
2020 Collaborative Motion Prediction via Neural Motion Message Passing
abstract
Motion prediction is essential and challenging for autonomous vehicles and social robots. One challenge of motion prediction is to model the interaction among traffic actors, which could cooperate with each other to avoid collisions or form groups. To address this challenge, we propose neural motion message passing (NMMP) to explicitly model the interaction and learn representations for directed interactions between actors. Based on the proposed NMMP, we design the motion prediction systems for two settings: the pedestrian setting and the joint pedestrian and vehicle setting. Both systems share a common pattern: we use an individual branch to model the behavior of a single actor and an interactive branch to model the interaction between actors, while with different wrappers to handle the varied input formats and characteristics. The experimental results show that both systems outperform the previous state-of-the-art methods on several existing benchmarks. Besides, we provide interpretability for interaction learning.
Yue Hu 0011, Siheng Chen, Ya Zhang 0002, Xiao Gu 0001
CVPR3
2020 Dynamic Multiscale Graph Neural Networks for 3D Skeleton Based Human Motion Prediction
abstract
We propose novel dynamic multiscale graph neural networks (DMGNN) to predict 3D skeleton-based human motions. The core idea of DMGNN is to use a multiscale graph to comprehensively model the internal relations of a human body for motion feature learning. This multiscale graph is adaptive during training and dynamic across network layers. Based on this graph, we propose a multiscale graph computational unit (MGCU) to extract features at individual scales and fuse features across scales. The entire model is action-category-agnostic and follows an encoder-decoder framework. The encoder consists of a sequence of MGCUs to learn motion features. The decoder uses a proposed graph-based gate recurrent unit to generate future poses. Extensive experiments show that the proposed DMGNN outperforms state-of-the-art methods in both short and long-term predictions on the datasets of Human 3.6M and CMU Mocap. We further investigate the learned multiscale graphs for the interpretability. The codes could be downloaded from https://github.com/limaosen0/DMGNN.
Maosen Li, Siheng Chen, Yangheng Zhao, Ya Zhang 0002, Yanfeng Wang 0001, Qi Tian 0001
CVPR4
2020 Iteratively-Refined Interactive 3D Medical Image Segmentation With Multi-Agent Reinforcement Learning
abstract
Existing automatic 3D image segmentation methods usually fail to meet the clinic use. Many studies have explored an interactive strategy to improve the image segmentation performance by iteratively incorporating user hints. However, the dynamic process for successive interactions is largely ignored. We here propose to model the dynamic process of iterative interactive image segmentation as a Markov decision process (MDP) and solve it with reinforcement learning (RL). Unfortunately, it is intractable to use single-agent RL for voxel-wise prediction due to the large exploration space. To reduce the exploration space to a tractable size, we treat each voxel as an agent with a shared voxel-level behavior strategy so that it can be solved with multi-agent reinforcement learning. An additional advantage of this multi-agent model is to capture the dependency among voxels for segmentation task. Meanwhile, to enrich the information of previous segmentations, we reserve the prediction uncertainty in the state space of MDP and derive an adjustment action space leading to a more precise and finer segmentation. In addition, to improve the efficiency of exploration, we design a relative cross-entropy gain-based reward to update the policy in a constrained direction. Experimental results on various medical datasets have shown that our method significantly outperforms existing state-of-the-art methods, with the advantage of less interactions and a faster convergence.
Xuan Liao, Wenhao Li 0001, Qisen Xu, Xiangfeng Wang 0001, Bo Jin 0003, Xiaoyun Zhang 0001, Yanfeng Wang 0001, Ya Zhang 0002
CVPR8
2020 FTL: A Universal Framework for Training Low-Bit DNNs via Feature Transfer
Kunyuan Du, Ya Zhang 0002, Haibing Guan, Qi Tian 0001, Yanfeng Wang 0001, Shenggan Cheng, James Lin 0001
ECCV (25)2
2020 Bottom-Up Temporal Action Localization with Mutual Regularization
Peisen Zhao, Lingxi Xie, Chen Ju, Ya Zhang 0002, Yanfeng Wang 0001, Qi Tian 0001
ECCV (8)4
2020 Knowledge Distillation from Multi-modal to Mono-modal Segmentation Networks
Minhao Hu, Matthis Maillard, Ya Zhang 0002, Tommaso Ciceri, Giammarco La Barbera, Isabelle Bloch, Pietro Gori
MICCAI (1)3
2020 Dual-Task Self-supervision for Cross-modality Domain Adaptation
Yingying Xue, Shixiang Feng, Ya Zhang 0002, Xiaoyun Zhang 0001, Yanfeng Wang 0001
MICCAI (1)3
2020 Graph Cross Networks with Vertex Infomax Pooling
abstract
We propose a novel graph cross network (GXN) to achieve comprehensive feature learning from multiple scales of a graph. Based on trainable hierarchical representations of a graph, GXN enables the interchange of intermediate features across scales to promote information flow. Two key ingredients of GXN include a novel vertex infomax pooling (VIPool), which creates multiscale graphs in a trainable manner, and a novel feature-crossing layer, enabling feature interchange across scales. The proposed VIPool selects the most informative subset of vertices based on the neural estimation of mutual information between vertex features and neighborhood features. The intuition behind is that a vertex is informative when it can maximally reflect its neighboring information. The proposed feature-crossing layer fuses intermediate features between two scales for mutual enhancement by improving information flow and enriching multiscale features at hidden layers. The cross shape of feature-crossing layer distinguishes GXN from many other multiscale architectures. Experimental results show that the proposed GXN improves the classification accuracy by 2.12% and 1.15% on average for graph classification and vertex classification, respectively. Based on the same network, the proposed VIPool consistently outperforms other graph-pooling methods.
Maosen Li, Siheng Chen, Ya Zhang 0002, Ivor W. Tsang
NeurIPS3
2020 A Unified Framework for Generalizable Style Transfer: Style and Content Separation
abstract
Image style transfer has drawn broad attention recently. However, most existing methods aim to explicitly model the transformation between different styles, and the learned model is often not generalizable to new styles. Based on the idea of style and content separation, we here propose a unified style transfer framework that consists of style encoder, content encoder, mixer and decoder. The style encoder and the content encoder are used to extract the style and content representations from the corresponding reference images. The two representations are integrated by the mixer and fed to the decoder, which generates images with the target style and content. Assuming the same encoder could be shared among different styles/contents, the style/content encoder explores a generalizable way to represent style/content information, i.e. the encoders are expected to capture the underlying representation for different styles/contents and generalize to new styles/contents. Training simultaneously with a number of styles and contents, the framework enables building one single transfer network for multiple styles and further leads to a key merit of the framework, i.e. its generalizability to new styles and contents. To evaluate the proposed framework, we apply it to both supervised and unsupervised style transfer, using character typeface transfer and neural style transfer as respective examples. For character typeface transfer, to separate the style features and content features, we leverage the conditional dependence of styles and contents given an image. For neural style transfer, we leverage the statistical information of feature maps in certain layers to represent style. Extensive experimental results have demonstrated the effectiveness and robustness of the proposed methods. Furthermore, models learned under the proposed framework are shown to be better generalizable to new styles and contents.
Yexun Zhang, Ya Zhang 0002, Wenbin Cai
IEEE Trans. Image Process.2
2019 Safeguarded Dynamic Label Regression for Noisy Supervision
abstract
Learning with noisy labels is imperative in the Big Data era since it reduces expensive labor on accurate annotations. Previous method, learning with noise transition, has enjoyed theoretical guarantees when it is applied to the scenario with the class-conditional noise. However, this approach critically depends on an accurate pre-estimated noise transition, which is usually impractical. Subsequent improvement adapts the preestimation in the form of a Softmax layer along with the training progress. However, the parameters in the Softmax layer are highly tweaked for the fragile performance and easily get stuck into undesired local minimums. To overcome this issue, we propose a Latent Class-Conditional Noise model (LCCN) that models the noise transition in a Bayesian form. By projecting the noise transition into a Dirichlet-distributed space, the learning is constrained on a simplex instead of some adhoc parametric space. Furthermore, we specially deduce a dynamic label regression method for LCCN to iteratively infer the latent true labels and jointly train the classifier and model the noise. Our approach theoretically safeguards the bounded update of the noise transition, which avoids arbitrarily tuning via a batch of samples. Extensive experiments have been conducted on controllable noise data with CIFAR10 and CIFAR-100 datasets, and the agnostic noise data with Clothing1M and WebVision17 datasets. Experimental results have demonstrated that the proposed model outperforms several state-of-the-art methods.
Jiangchao Yao, Hao Wu 0075, Ya Zhang 0002, Ivor W. Tsang, Jun Sun 0005
AAAI3
2019 Understanding VAEs in Fisher-Shannon Plane
abstract
In information theory, Fisher information and Shannon information (entropy) are respectively used to quantify the uncertainty associated with the distribution modeling and the uncertainty in specifying the outcome of given variables. These two quantities are complementary and are jointly applied to information behavior analysis in most cases. The uncertainty property in information asserts a fundamental trade-off between Fisher information and Shannon information, which enlightens us the relationship between the encoder and the decoder in variational auto-encoders (VAEs). In this paper, we investigate VAEs in the Fisher-Shannon plane, and demonstrate that the representation learning and the log-likelihood estimation are intrinsically related to these two information quantities. Through extensive qualitative and quantitative experiments, we provide with a better comprehension of VAEs in tasks such as high-resolution reconstruction, and representation learning in the perspective of Fisher information and Shannon information. We further propose a variant of VAEs, termed as Fisher auto-encoder (FAE), for practical needs to balance Fisher information and Shannon information. Our experimental results have demonstrated its promise in improving the reconstruction accuracy and avoiding the noninformative latent code as occurred in previous works.
Huangjie Zheng, Jiangchao Yao, Ya Zhang 0002, Ivor W. Tsang, Jia Wang 0004
AAAI3
2019 Actional-Structural Graph Convolutional Networks for Skeleton-Based Action Recognition
abstract
Action recognition with skeleton data has recently attracted much attention in computer vision. Previous studies are mostly based on fixed skeleton graphs, only capturing local physical dependencies among joints, which may miss implicit joint correlations. To capture richer dependencies, we introduce an encoder-decoder structure, called A-link inference module, to capture action-specific latent dependencies, i.e. actional links, directly from actions. We also extend the existing skeleton graphs to represent higher-order dependencies, i.e. structural links. Combing the two types of links into a generalized skeleton graph, We further propose the actional-structural graph convolution network (AS-GCN), which stacks actional-structural graph convolution and temporal convolution as a basic building block, to learn both spatial and temporal features for action recognition. A future pose prediction head is added in parallel to the recognition head to help capture more detailed action patterns through self-supervision. We validate AS-GCN in action recognition using two skeleton data sets, NTU-RGB+D and Kinetics. The proposed AS-GCN achieves consistently large improvement compared to the state-of-the-art methods. As a side product, AS-GCN also shows promising results for future pose prediction.
Maosen Li, Siheng Chen, Xu Chen 0026, Ya Zhang 0002, Yanfeng Wang 0001, Qi Tian 0001
CVPR4
2019 Online Popularity Prediction of Video Segments: Towards More Efficient Content Delivery Networks
abstract
Current state-of-the-art Online Video Services (OVSs) need to simultaneously serve hundreds of millions of users at peak time. To improve service efficiency, the caching strategies of their Content Delivery Networks (CDNs) usually work with video segments of a couple of seconds in length, which poses great challenges in popular prediction of video segments. Popularity prediction of entire videos is generally based on the correlation of early views and future views. However, for the popularity prediction of video segments, the first and the rest segments of a video should be distinguished first because users usually start watching a video from its first segment, but could skip any of the rest ones. Towards this end, we propose a novel method for video segment popularity prediction, in which, the popularity of videos' first and rest segments are predicted with different models. For the first segments, the popularity is predicted like entire videos by using a Multi-Linear Regression (MLR) Model based on past popularity. And for the rest segments, as the jumping viewing behavior has largely weaken the indicating power of their early views on future views, we predict their popularity based on the recent viewed previous segments. Specially, considering the real-time requirement of CDNs caching, it is infeasible to investigate the viewed segments sequences of users with sequence models, hence the prediction model is built on statistic level. Moreover, to maintain the performance in online prediction, we equip the two prediction models with online learning capability so as to timely capture the changes in popularity. For promptly service, the job is fulfilled with an online compensator instead of retraining the two prediction models with revealed target of past predictions. The online compensator estimates the prediction errors of the two models through a LSTM network based on their recent errors and the popularity data corresponded to recent tasks. Through extensive experiments on real-world data of iQIYI, one of the most popular OVS providers in China, we have demonstrated the promise of the proposed method.
Zhiyi Tan 0003, Ya Zhang 0002
GLOBECOM3
2019 Accelerate CNN via Recursive Bayesian Pruning
abstract
Channel Pruning, widely used for accelerating Convolutional Neural Networks, is an NP-hard problem due to the inter-layer dependency of channel redundancy. Existing methods generally ignored the above dependency for computation simplicity. To solve the problem, under the Bayesian framework, we here propose a layer-wise Recursive Bayesian Pruning method (RBP). A new dropout-based measurement of redundancy, which facilitate the computation of posterior assuming inter-layer dependency, is introduced. Specifically, we model the noise across layers as a Markov chain and target its posterior to reflect the inter-layer dependency. Considering the closed form solution for posterior is intractable, we derive a sparsity-inducing Dirac-like prior which regularizes the distribution of the designed noise to automatically approximate the posterior. Compared with the existing methods, no additional overhead is required when the inter-layer dependency assumed. The redundant channels can be simply identified by tiny dropout noise and directly pruned layer by layer. Experiments on popular CNN architectures have shown that the proposed method outperforms several state-of-the-arts. Particularly, we achieve up to 5.0x, 2.2x and 1.7x FLOPs reduction with little accuracy loss on the large scale dataset ILSVRC2012 for VGG16, ResNet50 and MobileNetV2, respectively.
Yuefu Zhou, Ya Zhang 0002, Yanfeng Wang 0001, Qi Tian 0001
ICCV2
2019 Collaborative Label Correction via Entropy Thresholding
abstract
Deep neural networks (DNNs) have the capacity to fit extremely noisy labels nonetheless they tend to learn data with clean labels first and then memorize those with noisy labels. We examine this behavior in light of the Shannon entropy of the predictions and demonstrate the low entropy predictions determined by a given threshold are much more reliable as the supervision than the original noisy labels. It also shows the advantage in maintaining more training samples than previous methods. Then, we power this entropy criterion with the Collaborative Label Correction (CLC) framework to further avoid undesired local minimums of the single network. A range of experiments have been conducted on multiple benchmarks with both synthetic and real-world settings. Extensive results indicate that our CLC outperforms several state-of-the-art methods.
Hao Wu 0075, Jiangchao Yao, Yinru Chen, Ya Zhang 0002, Yanfeng Wang 0001
ICDM5
2019 Deep Learning From Noisy Image Labels With Quality Embedding
abstract
There is an emerging trend to leverage noisy image datasets in many visual recognition tasks. However, the label noise among datasets severely degenerates the performance of deep learning approaches. Recently, one mainstream is to introduce the latent label to handle label noise, which has shown promising improvement in the network designs. Nevertheless, the mismatch between latent labels and noisy labels still affects the predictions in such methods. To address this issue, we propose a probabilistic model, which explicitly introduces an extra variable to represent the trustworthiness of noisy labels, termed as the quality variable. Our key idea is to identify the mismatch between the latent and noisy labels by embedding the quality variables into different subspaces, which effectively minimizes the influence of label noise. At the same time, reliable labels are still able to be applied for training. To instantiate the model, we further propose a Contrastive-Additive Noise network (CAN), which consists of two important layers: (1) the contrastive layer that estimates the quality variable in the embedding space to reduce the influence of noisy labels; and (2) the additive layer that aggregates the prior prediction and noisy labels as the posterior to train the classifier. Moreover, to tackle the challenges in optimization, we deduce an SGD algorithm with the reparameterization tricks, which makes our method scalable to big data.We validate the proposed method on a range of noisy image datasets. Comprehensive results have demonstrated that CAN outperforms the state-of-the-art deep learning approaches.
Jiangchao Yao, Ivor W. Tsang, Ya Zhang 0002, Jun Sun 0005, Chengqi Zhang
IEEE Trans. Image Process.4
2019 Predicting the Top-N Popular Videos via a Cross-Domain Hybrid Model
abstract
Predicting the top-N popular videos and their future views for a large batch of newly uploaded videos is of great commercial value to online video services (OVSs). Although many attempts have been made on video popularity prediction, the existing models has a much lower performance in predicting the top-N popular videos than that of the entire video set. The reason for this phenomenon is that most videos in an OVS system are unpopular, so models preferentially learn the popularity trends of unpopular videos to improve their performance on the entire video set. However, in most cases, it is critical to predict the performance on the top-N popular videos, which is the focus of this study. The challenge for the task are as follows. First, popular and unpopular videos may have similar early view patterns. Second, prediction models that are overly dependent on early view patterns limit the effects of other features. To address these challenges, we propose a novel multifactor differential influence prediction model based on multivariate linear regression. The model is designed to improve the discovery of popular videos and their popularity trends are learnt by enhancing the discriminative power of early patterns for different popularity trends and by optimizing the utilization of multisource data. We evaluate the proposed model using real-world YouTube data, and extensive experiments have demonstrated the effectiveness of our model.
Zhiyi Tan 0003, Ya Zhang 0002
IEEE Trans. Multim.2
2018 Chinese Handwriting Imitation with Hierarchical Generative Adversarial Network
Yujun Gu, Ya Zhang 0002, Yanfeng Wang 0001
BMVC3
2018 Learning Multi-touch Conversion Attribution with Dual-attention Mechanisms for Online Advertising
abstract
In online advertising, the Internet users may be exposed to a sequence of different ad campaigns, i.e., display ads, search, or referrals from multiple channels, before led up to any final sales conversion and transaction. For both campaigners and publishers, it is fundamentally critical to estimate the contribution from ad campaign touch-points during the customer journey (conversion funnel) and assign the right credit to the right ad exposure accordingly. However, the existing research on the multi-touch attribution problem lacks a principled way of utilizing the users' pre-conversion actions (i.e., clicks), and quite often fails to model the sequential patterns among the touch points from a user's behavior data. To make it worse, the current industry practice is merely employing a set of arbitrary rules as the attribution model, e.g., the popular last-touch model assigns 100% credit to the final touch-point regardless of actual attributions. In this paper, we propose a Dual-attention Recurrent Neural Network (DARNN) for the multi-touch attribution problem. It learns the attribution values through an attention mechanism directly from the conversion estimation objective. To achieve this, we utilize sequence-to-sequence prediction for user clicks, and combine both post-view and post-click attribution patterns together for the final conversion estimation. To quantitatively benchmark attribution models, we also propose a novel yet practical attribution evaluation scheme through the proxy of budget allocation (under the estimated attributions) over ad channels. The experimental results on two real datasets demonstrate the significant performance gains of our attribution model against the state of the art.
Kan Ren, Weinan Zhang 0001, Shuhao Liu 0002, Ya Zhang 0002, Yong Yu 0001, Jun Wang 0012
CIKM6
2018 Separating Style and Content for Generalized Style Transfer
abstract
Neural style transfer has drawn broad attention in recent years. However, most existing methods aim to explicitly model the transformation between different styles, and the learned model is thus not generalizable to new styles. We here attempt to separate the representations for styles and contents, and propose a generalized style transfer network consisting of style encoder, content encoder, mixer and decoder. The style encoder and content encoder are used to extract the style and content factors from the style reference images and content reference images, respectively. The mixer employs a bilinear model to integrate the above two factors and finally feeds it into a decoder to generate images with target style and content. To separate the style features and content features, we leverage the conditional dependence of styles and contents given an image. During training, the encoder network learns to extract styles and contents from two sets of reference images in limited size, one with shared style and the other with shared content. This learning framework allows simultaneous style transfer among multiple styles and can be deemed as a special 'multi-task' learning scenario. The encoders are expected to capture the underlying features for different styles and contents which is generalizable to new styles and contents. For validation, we applied the proposed algorithm to the Chinese Typeface transfer problem. Extensive experiment results on character generation have demonstrated the effectiveness and robustness of our method.
Yexun Zhang, Ya Zhang 0002, Wenbin Cai
CVPR2
2018 Multi-scale Spatially-Asymmetric Recalibration for Image Classification
Yan Wang 0033, Lingxi Xie, Siyuan Qiao, Ya Zhang 0002, Wenjun Zhang 0001, Alan L. Yuille
ECCV (13)4
2018 Unsupervised Local Facial Attributes Transfer Using Dual Discriminative Adversarial Networks
abstract
Facial attributes transfer is a challenging problem in computer vision and image processing. In this paper, we propose a new structure for local facial attributes transfer, called Dual Discriminative Adversarial Networks (DDAN). Given a source image with a certain local attribute (e.g. mouth open or wearing glasses), the DDAN model transforms source images with a certain attribute to target images without the same attribute while keeping the other attributes unchanged. A two-branch discriminator is introduced in the proposed DDAN model to simultaneously represent local and global features. Experimental results on the CelebA dataset have shown that the DDAN model outperforms several existing methods not only in terms of the quality of the generated images but also in terms of the convergence speed during model training.
Maosen Li, Ya Zhang 0002
ICME3
2018 Collaborative Learning for Weakly Supervised Object Detection
abstract
Weakly supervised object detection has recently received much attention, since it only requires image-level labels instead of the bounding-box labels consumed in strongly supervised learning. Nevertheless, the save in labeling expense is usually at the cost of model accuracy.In this paper, we propose a simple but effective weakly supervised collaborative learning framework to resolve this problem, which trains a weakly supervised learner and a strongly supervised learner jointly by enforcing partial feature sharing and prediction consistency. For object detection, taking WSDDN-like architecture as weakly supervised detector sub-network and Faster-RCNN-like architecture as strongly supervised detector sub-network, we propose an end-to-end Weakly Supervised Collaborative Detection Network. As there is no strong supervision available to train the Faster-RCNN-like sub-network, a new prediction consistency loss is defined to enforce consistency of predictions between the two sub-networks as well as within the Faster-RCNN-like sub-networks. At the same time, the two detectors are designed to partially share features to further guarantee the model consistency at perceptual level. Extensive experiments on PASCAL VOC 2007 and 2012 data sets have demonstrated the effectiveness of the proposed framework.
Jiangchao Yao, Ya Zhang 0002
IJCAI3
2018 Masking: A New Perspective of Noisy Supervision
abstract
It is important to learn various types of classifiers given training data with noisy labels. Noisy labels, in the most popular noise model hitherto, are corrupted from ground-truth labels by an unknown noise transition matrix. Thus, by estimating this matrix, classifiers can escape from overfitting those noisy labels. However, such estimation is practically difficult, due to either the indirect nature of two-step approaches, or not big enough data to afford end-to-end approaches. In this paper, we propose a human-assisted approach called ''Masking'' that conveys human cognition of invalid class transitions and naturally speculates the structure of the noise transition matrix. To this end, we derive a structure-aware probabilistic model incorporating a structure prior, and solve the challenges from structure extraction and structure alignment. Thanks to Masking, we only estimate unmasked noise transition probabilities and the burden of estimation is tremendously reduced. We conduct extensive experiments on CIFAR-10 and CIFAR-100 with three noise structures as well as the industrial-level Clothing1M with agnostic noise structure, and the results show that Masking can improve the robustness of classifiers significantly.
Bo Han 0003, Jiangchao Yao, Gang Niu 0001, Mingyuan Zhou, Ivor W. Tsang, Ya Zhang 0002, Masashi Sugiyama
NeurIPS6
2018 Deep Dual-view Network with Smooth Loss for Spinal Metastases Classification
abstract
Spinal metastases have a high incidence among cancer patients and may later develop to metastatic spinal cord compression (MSCC). Early detection of spinal metastases is critical for optimal treatment. The diagnosis is usually facilitated with computed tomography (CT) scans, which requires considerable efforts from well-trained radiologist. In this paper, we explore automatic spinal metastases classification based on CT images. Considering the unique characteristics of spinal CT images, a novel Deep Dual-view Network is proposed which contains two branches: a X-Y Conv Branch to extract the features for each individual cross-sectional image slice, and a Z Conv Branch to capture z-direction features from neighboring images. The features from the two branches are then fused to generate the final prediction, which imitates the doctor's way of combining cross sections with sagittal or coronal sections. Considering a tumor usually presents in multiple consecutive image slices, a smooth loss is introduced to maintain the label consistency of adjacent images. To validate the proposed approach, we collect a dataset of 316 patients with spinal metastases. Experimental results on this data set have demonstrated the effectiveness and the robustness of the proposed Deep Dual-view network.
Haoyan Guan, Guangyu Yao, Yexun Zhang, Yujun Gu, Ya Zhang 0002, Xiao Gu 0001
VCIP6
2018 Webly-Supervised Fine-Grained Visual Categorization via Deep Domain Adaptation
abstract
Learning visual representations from web data has recently attracted attention for object recognition. Previous studies have mainly focused on overcoming label noise and data bias and have shown promising results by learning directly from web data. However, we argue that it might be better to transfer knowledge from existing human labeling resources to improve performance at nearly no additional cost. In this paper, we propose a new semi-supervised method for learning via web data. Our method has the unique design of exploiting strong supervision, i.e., in addition to standard image-level labels, our method also utilizes detailed annotations including object bounding boxes and part landmarks. By transferring as much knowledge as possible from existing strongly supervised datasets to weakly supervised web images, our method can benefit from sophisticated object recognition algorithms and overcome several typical problems found in webly-supervised learning. We consider the problem of fine-grained visual categorization, in which existing training resources are scarce, as our main research objective. Comprehensive experimentation and extensive analysis demonstrate encouraging performance of the proposed approach, which, at the same time, delivers a new pipeline for fine-grained visual categorization that is likely to be highly effective for real-world applications.
Zhe Xu 0003, Shaoli Huang, Ya Zhang 0002, Dacheng Tao
IEEE Trans. Pattern Anal. Mach. Intell.3
2018 Stopping Criterion for Active Learning with Model Stability
abstract
Active learning selectively labels the most informative instances, aiming to reduce the cost of data annotation. While much effort has been devoted to active sampling functions, relatively limited attention has been paid to when the learning process should stop. In this article, we focus on the stopping criterion of active learning and propose a model stability--based criterion, that is, when a model does not change with inclusion of additional training instances. The challenge lies in how to measure the model change without labeling additional instances and training new models. Inspired by the stochastic gradient update rule, we use the gradient of the loss function at each candidate example to measure its effect on model change. We propose to stop active learning when the model change brought by any of the remaining unlabeled examples is lower than a given threshold. We apply the proposed stopping criterion to two popular classifiers: logistic regression (LR) and support vector machines (SVMs). In addition, we theoretically analyze the stability and generalization ability of the model obtained by our stopping criterion. Substantial experiments on various UCI benchmark datasets and ImageNet datasets have demonstrated that the proposed approach is highly effective.
Yexun Zhang, Wenbin Cai, Wenquan Wang, Ya Zhang 0002
ACM Trans. Intell. Syst. Technol.4
2018 Query-Free Clothing Retrieval via Implicit Relevance Feedback
abstract
Image-based clothing retrieval is receiving increasing interest with the growth of online shopping. In practice, users may often have a desired piece of clothing in mind (e.g., either having seen it before on the street or requiring certain specific clothing attributes), but may be unable to supply an image as a query. We model this problem as a new type of image retrieval task, in which the target image resides only in the user's mind (called “mental image retrieval” hereafter). Because of the absence of an explicit query image, we propose solving this problem through relevance feedback. Specifically, a new Bayesian formulation is proposed that simultaneously models the retrieval target and its high-level representation in the mind of the user (called the “user metric” hereafter) as posterior distributions of prefetched shop images and heterogeneous features extracted from multiple clothing attributes, respectively. Requiring only clicks as user feedback, the proposed algorithm is able to account for the variability in human decision-making. Experiments with real users demonstrate the effectiveness of the proposed algorithm.
Zhuoxiang Chen, Zhe Xu 0003, Ya Zhang 0002, Xiao Gu 0001
IEEE Trans. Multim.3
2018 Joint Latent Dirichlet Allocation for Social Tags
abstract
Social tags, serving as a textual source of simple but useful semantic metadata to reflect the user preference or describe the web objects, has been widely used in many applications. However, social tags have several unique characteristics, i.e., sparseness and data coupling (i.e., non-IIDness), which makes existing text analysis methods such as LDA not directly applicable. In this paper, we propose a new generative algorithm for social tag analysis named joint latent Dirichlet allocation, which models the generation of tags based on both the users and the objects, and thus accounts for the coupling relationships among social tags. The model introduces two latent factors that jointly influence tag generation: the user's latent interest factor and the object's latent topic factor, formulated as user-topic distribution matrix and object-topic distribution matrix, respectively. A Gibbs sampling approach is adopted to simultaneously infer the above two matrices as well as a topic-word distribution matrix. Experimental results on four social tagging datasets have shown that our model is able to capture more reasonable topics and achieves better performance than five state-of-the-art topic models in terms of the widely used point-wise mutual information metric. In addition, we analyze the learnt topics showing that our model recovers more themes from social tags while LDA may lead the topic vanishing problems, and demonstrate its advantages in the social recommendation by evaluating the retrieval results with mean reciprocal rank metric. Finally, we explore the joint procedure of our model in depth to show the non-IID characteristic of social tagging process.
Jiangchao Yao, Yanfeng Wang 0001, Ya Zhang 0002, Jun Sun 0005, Jun Zhou 0007
IEEE Trans. Multim.3
2017 SORT: Second-Order Response Transform for Visual Recognition
abstract
In this paper, we reveal the importance and benefits of introducing second-order operations into deep neural networks. We propose a novel approach named Second-Order Response Transform (SORT), which appends element-wise product transform to the linear sum of a two-branch network module. A direct advantage of SORT is to facilitate cross-branch response propagation, so that each branch can update its weights based on the current status of the other branch. Moreover, SORT augments the family of transform operations and increases the nonlinearity of the network, making it possible to learn flexible functions to fit the complicated distribution of feature space. SORT can be applied to a wide range of network architectures, including a branched variant of a chain-styled network and a residual network, with very light-weighted modifications. We observe consistent accuracy gain on both small (CIFAR10, CIFAR100 and SVHN) and big (ILSVRC2012) datasets. In addition, SORT is very efficient, as the extra computation overhead is less than 5%.
Yan Wang 0033, Lingxi Xie, Chenxi Liu 0001, Siyuan Qiao, Ya Zhang 0002, Wenjun Zhang 0001, Qi Tian 0001, Alan L. Yuille
ICCV5
2017 Distance metric learning with eigenvalue fine tuning
abstract
Distance metric learning focuses on learning one global or multiple local distance functions to draw similar instances close to each other and push away dissimilar ones. Most existing work has to do matrix projection to learn distance functions. In this paper, we present a novel distance function learning model which is based on eigenvalue fine tuning. Our model not only is able to learn the global distance function but also can be easily adopted into local metric learning tasks. From the perspective of dimension reduction, the proposed model can measure how much information has been preserved after feature transformation. Moreover, we connect our model with principal components analysis to improve its performance by introducing the label information. Experimental results have demonstrated the effectiveness of the proposed method.
Wenquan Wang, Ya Zhang 0002, Jinglu Hu
IJCNN2
2017 Discovering User Interests from Social Images
Jiangchao Yao, Ya Zhang 0002, Ivor W. Tsang, Jun Sun 0005
MMM (2)2
2017 Describing Geographical Characteristics with Social Images
Huangjie Zheng, Jiangchao Yao, Ya Zhang 0002
MMM (1)3
2017 From Theory to Practice: Efficient Active Cost-sensitive Classification with Expected Error Reduction
abstract
In many classification tasks, the data distribution is imbalanced and different misclassifications involve different costs. In addition, the data collected are often lack in labels and it is expensive and tedious to label them manually. Motivated by these two problems, we propose a novel active cost-sensitive classification algorithm based on the Expected Error Reduction (EER) framework, aiming to selectively label examples which can directly optimize the expected misclassification costs. However, the native EER (N-EER) framework is inefficient and impractical due to the considerable requirement for model retraining. In this paper, we propose an efficient EER (E-EER) to overcome the inefficiency of N-EER with the application of cost-sensitive classification which is realized by incorporating the cost information into the expected loss calculation. We first present a formal formulation for EER, then the active cost-sensitive classification algorithm is derived. In order to achieve E-EER, we derive an efficient model update rule for logistic regression (LR) and cost-sensitive support vector machines (C-SVM), respectively, to avoid model retraining, which are employed as the base learners. Furthermore, we theoretically analyze the error bound of our algorithm to provide a guarantee for its generalization performance. Extensive experiments demonstrate the effectiveness and efficiency of our method.
Yexun Zhang, Yanfeng Wang 0001, Wenbin Cai, Ya Zhang 0002
SDM5
2017 Clothing retrieval with visual attention model
abstract
Clothing retrieval is a challenging problem in computer vision. With the advance of Convolutional Neural Networks (CNNs), the accuracy of clothing retrieval has been significantly improved. FashionNet [1], a recent study, proposes to employ a set of artificial features in the form of landmarks for clothing retrieval, which are shown to be helpful for retrieval. However, the landmark detection module is trained with strong supervision which requires considerable efforts to obtain. In this paper, we propose a self-learning Visual Attention Model (VAM) to extracts attention maps from clothing images. The VAM is further connected to a global network to form an end-to-end network structure through Impdrop connection which randomly Dropout on the feature maps with the probabilities given by the attention map. Extensive experiments on several widely used benchmark clothing retrieval data sets have demonstrated the promise of the proposed method. We also show that compared to the trivial Product connection, the Impdrop connection makes the network structure more robust when training sets of limited size are used.
Yujun Gu, Ya Zhang 0002, Jun Zhou 0007, Xiao Gu 0001
VCIP3
2017 Deep hashing with triplet quantization loss
abstract
With the explosive growth of image databases, deep hashing, which learns compact binary descriptors for images, has become critical for fast image retrieval. Many existing deep hashing methods leverage quantization loss, defined as distance between the features before and after quantization, to reduce the error from binarizing features. While minimizing the quantization loss guarantees that quantization has minimal effect on retrieval accuracy, it unfortunately significantly reduces the expressiveness of features even before the quantization. In this paper, we show that the above definition of quantization loss is too restricted and in fact not necessary for maintaining high retrieval accuracy. We therefore propose a new form of quantization loss measured in triplets. The core idea of the triplet quantization loss is to learn discriminative real-valued descriptors which lead to minimal loss on retrieval accuracy after quantization. Extensive experiments on two widely used benchmark data sets of different scales, CIFAR-10 and In-shop, demonstrate that the proposed method outperforms the state-of-the-art deep hashing methods. Moreover, we show that the compact binary descriptors obtained with triplet quantization loss lead to very small performance drop after quantization.
Yuefu Zhou, Shanshan Huang 0007, Ya Zhang 0002, Yanfeng Wang 0001
VCIP3
2017 Friend or Foe: Fine-Grained Categorization With Weak Supervision
abstract
Multi-instance learning (MIL) is widely acknowledged as a fundamental method to solve weakly supervised problems. While MIL is usually effective in standard weakly supervised object recognition tasks, in this paper, we investigate the applicability of MIL on an extreme case of weakly supervised learning on the task of fine-grained visual categorization, in which intra-class variance could be larger than inter-class due to the subtle differences between subordinate categories. For this challenging task, we propose a new method that generalizes the standard multi-instance learning framework, for which a novel multi-task co-localization algorithm is proposed to take advantage of the relationship among fine-grained categories and meanwhile performs as an effective initialization strategy for the non-convex multi-instance objective. The localization results also enable object-level domain-specific fine-tuning of deep neural networks, which significantly boosts the performance. Experimental results on three fine-grained datasets reveal the effectiveness of the proposed method, especially the importance of exploiting inter-class relationships between object categories in weakly supervised fine-grained recognition.
Zhe Xu 0003, Dacheng Tao, Shaoli Huang, Ya Zhang 0002
IEEE Trans. Image Process.4
2017 Batch Mode Active Learning for Regression With Expected Model Change
abstract
While active learning (AL) has been widely studied for classification problems, limited efforts have been done on AL for regression. In this paper, we introduce a new AL framework for regression, expected model change maximization (EMCM), which aims at choosing the unlabeled data instances that result in the maximum change of the current model once labeled. The model change is quantified as the difference between the current model parameters and the updated parameters after the inclusion of the newly selected examples. In light of the stochastic gradient descent learning rule, we approximate the change as the gradient of the loss function with respect to each single candidate instance. Under the EMCM framework, we propose novel AL algorithms for the linear and nonlinear regression models. In addition, by simulating the behavior of the sequential AL policy when applied for k iterations, we further extend the algorithms to batch mode AL to simultaneously choose a set of k most informative instances at each query time. Extensive experimental results on both UCI and StatLib benchmark data sets have demonstrated that the proposed algorithms are highly effective and efficient.
Wenbin Cai, Muhan Zhang, Ya Zhang 0002
IEEE Trans. Neural Networks Learn. Syst.3
2017 Active Learning for Classification with Maximum Model Change
abstract
Most existing active learning studies focus on designing sample selection algorithms. However, several fundamental problems deserve investigation to provide deep insight into active learning. In this article, we conduct an in-depth investigation on active learning for classification from the perspective of model change. We derive a general active learning framework for classification called maximum model change (MMC), which aims at querying the influential examples. The model change is quantified as the difference between the model parameters before and after training with the expanded training set. Inspired by the stochastic gradient update rule, the gradient of the loss with respect to a given candidate example is adopted to approximate the model change. This framework is applied to two popular classifiers: support vector machines and logistic regression. We analyze the convergence property of MMC and theoretically justify it. We explore the connection between MMC and uncertainty-based sampling to provide a uniform view. In addition, we discuss its potential usability to other learning models and show its applicability in a wide range of applications. We validate the MMC strategy on two kinds of benchmark datasets, the UCI repository and ImageNet, and show that it outperforms many state-of-the-art methods.
Wenbin Cai, Yexun Zhang, Ya Zhang 0002, Wenquan Wang, Zhuoxiang Chen, Chris Ding
ACM Trans. Inf. Syst.3
2016 Synergy and antagonism in online advertising
abstract
An advertising campaign is usually composed of a series of coordinated advertisements, with various formats and delivered through different media channels. Several existing studies have attempted to measure the individual contribution of related advertisements in a campaign, resulting in rule-based and data-driven multi-touch attribution models. However, most of these models ignored the interaction among different advertisements. We here specifically consider the synergy and antagonism interaction among advertisements and propose a data-driven model. With the interaction, the influence of different advertisements is no longer linearly additive. We choose a vector representation for the influence of an individual advertisement, which can easily model the synergy and antagonism interactions. We validate the proposed method with a real-world dataset and show that antagonism is prevalent among advertisements while synergy rarely appears. In addition, our method leads to more accurate purchase predictions.
Ya Zhang 0002, Xiao Gu 0001
BDCAT2
2016 Part-Stacked CNN for Fine-Grained Visual Categorization
abstract
In the context of fine-grained visual categorization, the ability to interpret models as human-understandable visual manuals is sometimes as important as achieving high classification accuracy. In this paper, we propose a novel Part-Stacked CNN architecture that explicitly explains the finegrained recognition process by modeling subtle differences from object parts. Based on manually-labeled strong part annotations, the proposed architecture consists of a fully convolutional network to locate multiple object parts and a two-stream classification network that encodes object-level and part-level cues simultaneously. By adopting a set of sharing strategies between the computation of multiple object parts, the proposed architecture is very efficient running at 20 frames/sec during inference. Experimental results on the CUB-200-2011 dataset reveal the effectiveness of the proposed architecture, from multiple perspectives of classification accuracy, model interpretability, and efficiency. Being able to provide interpretable recognition results in realtime, the proposed method is believed to be effective in practical applications.
Shaoli Huang, Zhe Xu 0003, Dacheng Tao, Ya Zhang 0002
CVPR4
2016 Preference Aware Recommendation Based on Categorical Information
abstract
Contextual aware matrix factorization has been widely used in recommender systems by learning latent feature vectors of users and items along with contextual information. While most of them add identical bias for each type of side information to represent systematic tendencies in users' rating behaviors, they are not able to capture the preference unique to users or items. In this paper, we propose a probabilistic generative model which allows the bias to vary among different types of users or items. We first use Gaussian Mixture Components to cluster the users (or items) based on corresponding latent feature vectors respectively. Biases are then distributed on these clusters along with categorical side information. Finally, they are jointed with latent feature vectors of the users and items to affect the generation of observed ratings. Experiments on MovieLens-100K and MovieLens-1M data sets have shown promising results compared with state-of-the-art contextual aware recommendation approaches. We also qualitatively analyze the preferences of users and items and demonstrate differences in preference among both users and items.
Zhiwei Rao, Jiangchao Yao, Ya Zhang 0002
ICMLA3
2016 A parameter partial-sharing CNN architecture for cross-domain clothing retrieval
abstract
Cross-domain clothing retrieval is a challenging task due to significant differences between online shop images taken in controlled conditions of clean backgrounds, good lighting, and fixed poses, and street photos captured in uncontrollable conditions. In recent years, Convolutional Neural Networks (CNNs) have demonstrated its effectiveness for various computer vision problems including image retrieval. There are two mainstream CNNs based models addressing image retrieval tasks: triplet network models [1] and siamese network models [2]. In this paper, we first make a thorough comparison between the two types of models, and investigate the impact of different domain adaptation schemes including parameter sharing, non-sharing, and a new partial-sharing strategy between the street domain and the shop domain. Extensive experiments have revealed that the proposed partial-sharing scheme is able to reduce the number of parameters by a significant margin, while achieving comparable retrieval accuracy as the state-of-the-art scheme using triplet loss with non-sharing parameters.
Yichao Xiong, Zhe Xu 0003, Ya Zhang 0002
VCIP4
2016 Multinomial Latent Logistic Regression for Image Understanding
abstract
In this paper, we present multinomial latent logistic regression (MLLR), a new learning paradigm that introduces latent variables to logistic regression. By inheriting the advantages of logistic regression, MLLR is efficiently optimized using the second-order derivatives and provides effective probabilistic analysis on output predictions. MLLR is particularly effective in weakly supervised settings where the latent variable has an exponential number of possible values. The effectiveness of MLLR is demonstrated on four different image understanding applications, including a new challenging architectural style classification task. Furthermore, we show that MLLR can be generalized to general structured output prediction, and in doing so, we provide a thorough investigation of the connections and differences between MLLR and existing related algorithms, including latent structural SVMs and hidden conditional random fields.
Zhe Xu 0003, Zhibin Hong, Ya Zhang 0002, Junjie Wu 0002, Ah Chung Tsoi, Dacheng Tao
IEEE Trans. Image Process.3
2015 IOHMM for location prediction with missing data
abstract
In recent years, the widespread adoption of GPS enabled vehicles brings the Location Based Services new opportunities. It benefits many related fields such as urban planning, city traffic modeling, personalized recommendations and driving suggestions. The service providers can understand their users better by modeling the mobility pattern and provide more personalized services by predicting the destination of users' travels. In this paper, we propose to model both the temporal and spatial mobility patterns of human movements and predict the user's travel destination from certain origin place at certain time with specific IOHMM. In order to account for data missing, we introduce a dummy state in the process of constructing the IOHMM data sequence. We also demonstrate the possibility to represent individual mobility preference by building the user mobility profiles with the learnt IOHMM. We evaluate the prediction accuracy of our method with two datasets, and the experimental results show that our method outperforms several state-of-the-art works on both datasets.
Yanfeng Wang 0001, Ya Zhang 0002
DSAA3
2015 Augmenting Strong Supervision Using Web Data for Fine-Grained Categorization
abstract
We propose a new method for fine-grained object recognition that employs part-level annotations and deep convolutional neural networks (CNNs) in a unified framework. Although both schemes have been widely used to boost recognition performance, due to the difficulty in acquiring detailed part annotations, strongly supervised fine-grained datasets are usually too small to keep pace with the rapid evolution of CNN architectures. In this paper, we solve this problem by exploiting inexhaustible web data. The proposed method improves classification accuracy in two ways: more discriminative CNN feature representations are generated using a training set augmented by collecting a large number of part patches from weakly supervised web images, and more robust object classifiers are learned using a multi-instance learning algorithm jointly on the strong and weak datasets. Despite its simplicity, the proposed method delivers a remarkable performance improvement on the CUB200-2011 dataset compared to baseline part-based R-CNN methods, and achieves the highest accuracy on this dataset even in the absence of test image annotations.
Zhe Xu 0003, Shaoli Huang, Ya Zhang 0002, Dacheng Tao
ICCV3
2015 Joint Latent Dirichlet Allocation for non-iid social tags
abstract
Topic models have been widely used for analyzing text corpora and achieved great success in applications including content organization and information retrieval. However, different from traditional text data, social tags in the web containers are usually of small amounts, unordered, and non-iid, i.e., it is highly dependent on contextual information such as users and objects. Considering the specific characteristics of social tags, we here introduce a new model named Joint Latent Dirichlet Allocation (JLDA) to capture the relationships among users, objects, and tags. The model assumes that the latent topics of users and those of objects jointly influence the generation of tags. The latent distributions is then inferred with Gibbs sampling. Experiments on two social tag data sets have demonstrated that the model achieves a lower predictive error and generates more reasonable topics. We also present an interesting application of this model to object recommendation.
Jiangchao Yao, Ya Zhang 0002, Zhe Xu 0003, Jun Sun 0005, Jun Zhou 0007, Xiao Gu 0001
ICME2
2015 Online Learning Algorithm for Collective LDA
abstract
Collective Latent Dirichlet Allocation (C-LDA) is proposed as an extension of LDA to simultaneously model multiple corpora from different domains in order to overcome bias of individual corpus. However, with large volume of document collections from various sources, it becomes challenging to achieve fast convergence for C-LDA. The high time complexity of C-LDA limits its application to real-world tasks. Luckily, online learning has shown promise for speeding up the convergence of LDA. In this paper, we propose to explore online learning for collective LDA (OVCLDA). We first develop an efficient variational inference algorithm for collective LDA and then extend it to the online learning framework. We perform experiments with various real-world corpora. Experimental results have shown that OVCLDA can learn comparable topics with C-LDA and better than Online LDA, and achieves comparable computational efficiency with Online LDA and is much more efficient than C-LDA.
Jiangchao Yao, Yanfeng Wang 0001, Ya Zhang 0002
ICMLA4
2015 Active learning for ranking with sample density
Wenbin Cai, Muhan Zhang, Ya Zhang 0002
Inf. Retr. J.3
2015 Deep feature for text-dependent speaker verification
Yanmin Qian, Nanxin Chen, Tianfan Fu, Ya Zhang 0002, Kai Yu 0004
Speech Commun.5
2015 Active Learning for Ranking through Expected Loss Optimization
abstract
Learning to rank arises in many data mining applications, ranging from web search engine, online advertising to recommendation system. In learning to rank, the performance of a ranking model is strongly affected by the number of labeled examples in the training set; on the other hand, obtaining labeled examples for training data is very expensive and time-consuming. This presents a great need for the active learning approaches to select most informative examples for ranking learning; however, in the literature there is still very limited work to address active learning for ranking. In this paper, we propose a general active learning framework, expected loss optimization (ELO), for ranking. The ELO framework is applicable to a wide range of ranking functions. Under this framework, we derive a novel algorithm, expected discounted cumulative gain (DCG) loss optimization (ELO-DCG), to select most informative examples. Then, we investigate both query and document level active learning for raking and propose a two-stage ELO-DCG algorithm which incorporate both query and document selection into active learning. Furthermore, we show that it is flexible for the algorithm to deal with the skewed grade distribution problem with the modification of the loss function. Extensive experiments on real-world web search data sets have demonstrated great potential and effectiveness of the proposed framework and algorithms.
Bo Long, Jiang Bian 0002, Olivier Chapelle, Ya Zhang 0002, Yoshiyuki Inagaki, Yi Chang 0001
IEEE Trans. Knowl. Data Eng.4
2015 Active Learning for Web Search Ranking via Noise Injection
abstract
Learning to rank has become increasingly important for many information retrieval applications. To reduce the labeling cost at training data preparation, many active sampling algorithms have been proposed. In this article, we propose a novel active learning-for-ranking strategy called ranking-based sensitivity sampling (RSS), which is tailored for Gradient Boosting Decision Tree (GBDT), a machine-learned ranking method widely used in practice by major commercial search engines for ranking. We leverage the property of GBDT that samples close to the decision boundary tend to be sensitive to perturbations and design the active learning strategy accordingly. We further theoretically analyze the proposed strategy by exploring the connection between the sensitivity used for sample selection and model regularization to provide a potentially theoretical guarantee w.r.t. the generalization capability. Considering that the performance metrics of ranking overweight the top-ranked items, item rank is incorporated into the selection function. In addition, we generalize the proposed technique to several other base learners to show its potential applicability in a wide variety of applications. Substantial experimental results on both the benchmark dataset and a real-world dataset have demonstrated that our proposed active learning strategy is highly effective in selecting the most informative examples.
Wenbin Cai, Muhan Zhang, Ya Zhang 0002
ACM Trans. Web3
2014 Feature Selection at the Discrete Limit
abstract
Feature selection plays an important role in many machine learning and data mining applications. In this paper, we propose to use L2,p norm for feature selection with emphasis on small p. As p approaches 0, feature selection becomes discrete feature selection problem. We provide two algorithms, proximal gradient algorithm and rank one update algorithm, which is more efficient at large regularization. We provide closed form solutions of the proximal operator at p = 0, 1/2. Experiments onreal life datasets show that features selected at small p consistently outperform features selected at p = 1, the standard L2,1 approach and other popular feature selection methods.
Chris Ding, Ya Zhang 0002, Feiping Nie 0001
AAAI3
2014 Architectural Style Classification Using Multinomial Latent Logistic Regression
Zhe Xu 0003, Dacheng Tao, Ya Zhang 0002, Junjie Wu 0002, Ah Chung Tsoi
ECCV (1)3
2014 Stability-Based Stopping Criterion for Active Learning
abstract
While active learning has drawn broad attention in recent years, there are relatively few studies on stopping criterion for active learning. We here propose a novel model stability based stopping criterion, which considers the potential of each unlabeled examples to change the model once added to the training set. The underlying motivation is that active learning should terminate when the model does not change much by adding remaining examples. Inspired by the widely used stochastic gradient update rule, we use the gradient of the loss at each candidate example to measure its capability to change the classifier. Under the model change rule, we stop active learning when the changing ability of all remaining unlabeled examples is less than a given threshold. We apply the stability-based stopping criterion to two popular classifiers: logistic regression and support vector machines (SVMs). It can be generalized to a wide spectrum of learning models. Substantial experimental results on various UCI benchmark data sets have demonstrated that the proposed approach outperforms state-of-art methods in most cases.
Wenquan Wang, Wenbin Cai, Ya Zhang 0002
ICDM3
2014 Multi-touch Attribution in Online Advertising with Survival Theory
abstract
Multi-touch attribution, which allows distributing the credit to all related advertisements based on their corresponding contributions, has recently become an important research topic in digital advertising. Traditionally, rule-based attribution models have been used in practice. The drawback of such rule-based models lies in the fact that the rules are not derived form the data but only based on simple intuition. With the ever enhanced capability to tracking advertisement and users' interaction with the advertisement, data-driven multi-touch attribution models, which attempt to infer the contribution from user interaction data, become an important research direction. We here propose a new data-driven attribution model based on survival theory. By adopting a probabilistic framework, one key advantage of the proposed model is that it is able to remove the presentation biases inherit to most of the other attribution models. In addition to model the attribution, the proposed model is also able to predict user's 'conversion' probability. We validate the proposed method with a real-world data set obtained from a operational commercial advertising monitoring company. Experiment results have shown that the proposed method is quite promising in both conversion prediction and attribution.
Ya Zhang 0002, Jianbiao Ren
ICDM1
2014 Active Learning for Support Vector Machines with Maximum Model Change
Wenbin Cai, Ya Zhang 0002, Wenquan Wang, Chris Ding, Xiao Gu 0001
ECML/PKDD (1)2
2014 Social Image Analysis From a Non-IID Perspective
abstract
An image in social media, termed a social image, exhibits characteristics different from images widely discussed in image processing. They can be described by both content and social related attributes, called social image attributes, including visual contents, users, tags, and timestamps. There are strong coupling relationships between social image attributes, which make social images not independent and identically distributed (non-IID). By analyzing the relationships among these attributes, we can better understand the semantic activities conducted on such non-IID social images, hence enabling new applications including content organization, recommendation, and social activity understanding. In this article, we present a novel algorithm to analyze the coupling relationships between social images, which involves not only intra-coupled similarity within a social image attribute, but also inter-coupled similarity between attributes, in analyzing the non-IIDness of the similarity between social images. In particular, we propose a multi-entry version of the coupled similarity metric to deal with attributes (i.e., tags) which have a many-to-one relationship with respect to images. Experimental results on a Flickr group dataset show that the proposed algorithm captures coupling relationships and therefore achieves promising results in various applications, including image clustering and tagging.
Zhe Xu 0003, Ya Zhang 0002, Longbing Cao
IEEE Trans. Multim.2
2013 Social-Correlation Based Mutual Reinforcement for Short Text Classification and User Interest Tagging
Ya Zhang 0002
ADMA (1)2
2013 Joint Optimization for Consistent Multiple Graph Matching
abstract
The problem of graph matching in general is NP-hard and approaches have been proposed for its sub optimal solution, most focusing on finding the one-to-one node mapping between two graphs. A more general and challenging problem arises when one aims to find consistent mappings across a number of graphs more than two. Conventional graph pair matching methods often result in mapping inconsistency since the mapping between two graphs can either be determined by pair mapping or by an additional anchor graph. To address this issue, a novel formulation is derived which is maximized via alternating optimization. Our method enjoys several advantages: 1) the mappings are jointly optimized rather than sequentially performed by applying pair matching, allowing the global affinity information across graphs can be propagated and explored, 2) the number of concerned variables to optimize is in linear with the number of graphs, being superior to local pair matching resulting in O(n^2) variables, 3) the mapping consistency constraints are analytically satisfied during optimization, and 4) off-the-shelf graph pair matching solvers can be reused under the proposed framework in an `out-of-the-box' fashion. Competitive results on both the synthesized data and the real data are reported, by varying the level of deformation, outliers and edge densities.
Junchi Yan, Hongyuan Zha, Xiaokang Yang 0001, Ya Zhang 0002, Stephen M. Chu
ICCV5
2013 Maximizing Expected Model Change for Active Learning in Regression
abstract
Active learning is well-motivated in many supervised learning tasks where unlabeled data may be abundant but labeled examples are expensive to obtain. The goal of active learning is to maximize the performance of a learning model using as few labeled training data as possible, thereby minimizing the cost of data annotation. So far, there is still very limited work on active learning for regression. In this paper, we propose a new active learning framework for regression called Expected Model Change Maximization (EMCM), which aims to choose the examples that lead to the largest change to the current model. The model change is measured as the difference between the current model parameters and the updated parameters after training with the enlarged training set. Inspired by the Stochastic Gradient Descent (SGD) update rule, the change is estimated as the gradient of the loss with respect to a candidate example for active learning. Under this framework, we derive novel active learning algorithms for both linear regression and nonlinear regression to select the most informative examples. Extensive experimental results on the benchmark data sets from UCI machine learning repository have demonstrated that the proposed algorithms are highly effective in choosing the most informative examples and robust to various types of data distributions.
Wenbin Cai, Ya Zhang 0002, Jun Zhou 0007
ICDM2
2013 Face recogntion in open world environment
abstract
Face recognition in open world environment is a very challenging task due to variant appearances of the target persons and a large scale of unregistered probe faces. In this paper we combine two parallel classifiers, one based on the Local Binary Pattern (LBP) feature and the other based on the Gabor features, to build a specific face recognizer for each target person. Faces used for training are borderline patterns obtained through a morphing procedure combing target faces and random non-target ones. Grid-search is applied to find an optimal morphing-degree-pair. By using an AND operator to integrate the prediction of the two complementary parallel classifiers, many false positives are eliminated in the final results. The proposed algorithm is compared with the Robust Sparse Coding method, using selected celebrities as the target persons and the images from FERET as the non-target faces. Experimental results suggest that the proposed approach is better at tolerating the distortion of the target person's appearance and has a lower false alarm rate.
Jieqiong Qiu, Ya Zhang 0002, Jun Sun 0005
VCIP2
2013 Collaborative filtering with social regularization for TV program recommendation
Ya Zhang 0002, Weiyuan Chen, Zibin Yin
Knowl. Based Syst.1
2013 Reorder user's tweets
abstract
Twitter displays the tweets a user received in a reversed chronological order, which is not always the best choice. As Twitter is full of messages of very different qualities, many informative or relevant tweets might be flooded or displayed at the bottom while some nonsense buzzes might be ranked higher. In this work, we present a supervised learning method for personalized tweets reordering based on user interests. User activities on Twitter, in terms of tweeting, retweeting, and replying, are leveraged to obtain the training data for reordering models. Through exploring a rich set of social and personalized features, we model the relevance of tweets by minimizing the pairwise loss of relevant and irrelevant tweets. The tweets are then reordered according to the predicted relevance scores. Experimental results with real twitter user activities demonstrated the effectiveness of our method. The new method achieved above 30% accuracy gain compared with the default ordering in twitter based on time.
Keyi Shen, Jianmin Wu, Ya Zhang 0002, Yiping Han, Xiaokang Yang 0001, Li Song 0001, Xiao Gu 0001
ACM Trans. Intell. Syst. Technol.3
2012 Variance maximization via noise injection for active sampling in learning to rank
abstract
Active learning for ranking, which is to selectively label the most informative examples, has been widely studied in recent years. In this paper, we propose a general active learning for ranking strategy called Variance Maximization (VM). The algorithm relies on noise injection to perturb the original unlabeled examples and generate the rank distribution of each example. Using a DCG-like gain function to measure each ranked list sampled from the rank distribution, Variance Maximization selects the unlabeled example with the largest variance in the gain. The VM strategy is applied at both the query level and the document level, and a two-stage active learning algorithm is further derived. Experimental results on both the LETOR 4.0 dataset and a real-world Web search ranking dataset have demonstrated the effectiveness of the proposed active learning approach.
Wenbin Cai, Ya Zhang 0002
CIKM2
2012 On the Convergence of Graph Matching: Graduated Assignment Revisited
Junchi Yan, Hequan Zhang, Ya Zhang 0002, Xiaokang Yang 0001, Hongyuan Zha
ECCV (3)4
2011 Boosted multi-task learning
Olivier Chapelle, Pannagadatta K. Shivaswamy, Srinivas Vadrevu, Kilian Q. Weinberger, Ya Zhang 0002, Belle L. Tseng
Mach. Learn.5
2010 Multi-task learning for boosting with application to web search ranking
abstract
In this paper we propose a novel algorithm for multi-task learning with boosted decision trees. We learn several different learning tasks with a joint model, explicitly addressing the specifics of each learning task with task-specific parameters and the commonalities between them through shared parameters. This enables implicit data sharing and regularization. We evaluate our learning method on web-search ranking data sets from several countries. Here, multitask learning is particularly helpful as data sets from different countries vary largely in size because of the cost of editorial judgments. Our experiments validate that learning various tasks jointly can lead to significant improvements in performance with surprising reliability.
Olivier Chapelle, Pannagadatta K. Shivaswamy, Srinivas Vadrevu, Kilian Q. Weinberger, Ya Zhang 0002, Belle L. Tseng
KDD5
2010 Active learning for ranking through expected loss optimization
abstract
Learning to rank arises in many information retrieval applications, ranging from Web search engine, online advertising to recommendation system. In learning to rank, the performance of a ranking model is strongly affected by the number of labeled examples in the training set; on the other hand, obtaining labeled examples for training data is very expensive and time-consuming. This presents a great need for the active learning approaches to select most informative examples for ranking learning; however, in the literature there is still very limited work to address active learning for ranking. In this paper, we propose a general active learning framework, Expected Loss Optimization (ELO), for ranking. The ELO framework is applicable to a wide range of ranking functions. Under this framework, we derive a novel algorithm, Expected DCG Loss Optimization (ELO-DCG), to select most informative examples. Furthermore, we investigate both query and document level active learning for raking and propose a two-stage ELO-DCG algorithm which incorporate both query and document selection into active learning. Extensive experiments on real-world Web search data sets have demonstrated great potential and effective-ness of the proposed framework and algorithms.
Bo Long, Olivier Chapelle, Ya Zhang 0002, Yi Chang 0001, Zhaohui Zheng 0001, Belle L. Tseng
SIGIR3
2010 Learning more powerful test statistics for click-based retrieval evaluation
abstract
Interleaving experiments are an attractive methodology for evaluating \nretrieval functions through implicit feedback. Designed as \na blind and unbiased test for eliciting a preference between two \nretrieval functions, an interleaved ranking of the results of two retrieval \nfunctions is presented to the users. It is then observed whether \nthe users click more on results from one retrieval function or the \nother. While it was shown that such interleaving experiments reliably \nidentify the better of the two retrieval functions, the naive \napproach of counting all clicks equally leads to a suboptimal test. \nWe present new methods for learning how to score different types \nof clicks so that the resulting test statistic optimizes the statistical \npower of the experiment. This can lead to substantial savings in \nthe amount of data required for reaching a target confidence level. \nOur methods are evaluated on an operational search engine over a \ncollection of scientific articles.
Yisong Yue, Yue Gao 0005, Olivier Chapelle, Ya Zhang 0002, Thorsten Joachims
SIGIR4
2009 Expected reciprocal rank for graded relevance
abstract
While numerous metrics for information retrieval are available in the case of binary relevance, there is only one commonly used metric for graded relevance, namely the Discounted Cumulative Gain (DCG). A drawback of DCG is its additive nature and the underlying independence assumption: a document in a given position has always the same gain and discount independently of the documents shown above it. Inspired by the "cascade" user model, we present a new editorial metric for graded relevance which overcomes this difficulty and implicitly discounts documents which are shown below very relevant documents. More precisely, this new metric is defined as the expected reciprocal length of time that the user will take to find a relevant document. This can be seen as an extension of the classical reciprocal rank to the graded relevance case and we call this metric Expected Reciprocal Rank (ERR). We conduct an extensive evaluation on the query logs of a commercial search engine and show that ERR correlates better with clicks metrics than other editorial metrics.
Olivier Chapelle, Donald Metlzer, Ya Zhang 0002, Pierre Grinspan
CIKM3
2009 A risk minimization framework for domain adaptation
abstract
Supervised learning algorithms usually require high quality labeled training set of large volume. It is often expensive to obtain such labeled examples in every domain of an application. Domain adaptation aims to help in such cases by utilizing data available in related domains. However transferring knowledge from one domain to another is often non trivial due to different data distributions among the domains. Moreover, it is usually very hard to measure and formulate these distribution differences. Hence we introduce a new concept of label-relation function to transfer knowledge among different domains without explicitly formulating the data distribution differences. A novel learning framework, Domain Transfer Risk Minimization (DTRM), is proposed based on this concept. DTRM simultaneously minimizes the empirical risk for the target and the regularized empirical risk for source domain. Under this framework, we further derive a generic algorithm called Domain Adaptation by Label Relation (DALR) that is applicable to various applications in both classification and regression settings. DALR iteratively updates the target hypothesis function and outputs for the source domain until it converges. We provide an in-depth theoretical analysis of DTRM and establish fundamental error bounds. We also experimentally evaluate DALR on the task of ranking search results using real-world data. Our experimental results show that the proposed algorithm effectively and robustly utilizes data from source domains under various conditions: different sizes for source domain data; different noise levels for source domain data, and different difficulty levels for target domain data.
Bo Long, Sudarshan Lamkhede, Srinivas Vadrevu, Ya Zhang 0002, Belle L. Tseng
CIKM4
2009 A dynamic bayesian network click model for web search ranking
abstract
As with any application of machine learning, web search ranking requires labeled data. The labels usually come in the form of relevance assessments made by editors. Click logs can also provide an important source of implicit feedback and can be used as a cheap proxy for editorial labels. The main difficulty however comes from the so called position bias - urls appearing in lower positions are less likely to be clicked even if they are relevant. In this paper, we propose a Dynamic Bayesian Network which aims at providing us with unbiased estimation of the relevance from the click logs. Experiments show that the proposed click model outperforms other existing click models in predicting both click-through rate and relevance.
Olivier Chapelle, Ya Zhang 0002
WWW2
2008 A Two-Stage Approach to Chinese Part-of-Speech Tagging
Aitao Chen, Ya Zhang 0002, Gordon Sun
IJCNLP2
2008 Identifying regional sensitive queries in web search
abstract
In Web search ranking, the expected results for some queries could vary greatly depending upon location of the user. We name such queries regional sensitive queries. Identifying regional sensitivity of queries is important to meet users' needs. The objective of this work is to identify whether a user expects only regional results for a query. We present three novel features generated from search logs and build a meta query classifier to identify regional sensitive query. Experimental results show that the proposed method achieves high accuracy in identifying regional sensitive queries.
Srinivas Vadrevu, Ya Zhang 0002, Belle L. Tseng, Gordon Sun
WWW2
2007 Normalized Linear Transform for Cross-Platform Microarray Data Integration
abstract
With microarray data being dramatically accumulated, integrating data from related studies represents a natural way to increase sample size so that more reliable statistical analysis may be performed. However, inherent variation among different microarray platforms makes the data integration not a trivial task. In this paper, we present a simple and effective integration scheme, called normalized linear transform (NLT), to combine data from different microarray platforms. The NLT scheme is compared with three other integration schemes for two tasks: classification analysis and gene marker selection. Our experiments demonstrate that the NLT scheme performs best in terms of classification accuracy under various classification settings, and leads to more biologically significant marker genes.
Huilin Xiong, Ya Zhang 0002, Xue-wen Chen 0001
ICMLA2
2007 Data-Dependent Kernel Machines for Microarray Data Classification
abstract
One important application of gene expression analysis is to classify tissue samples according to their gene expression levels. Gene expression data are typically characterized by high dimensionality and small sample size, which makes the classification task quite challenging. In this paper, we present a data-dependent kernel for microarray data classification. This kernel function is engineered so that the class separability of the training data is maximized. A bootstrapping-based resampling scheme is introduced to reduce the possible training bias. The effectiveness of this adaptive kernel for microarray data classification is illustrated with a k-Nearest Neighbor (KNN) classifier. Our experimental study shows that the data-dependent kernel leads to a significant improvement in the accuracy of KNN classifiers. Furthermore, this kernel-based KNN scheme has been demonstrated to be competitive to, if not better than, more sophisticated classifiers such as Support Vector Machines (SVMs) and the Uncorrelated Linear Discriminant Analysis (ULDA) for classifying gene expression data.
Huilin Xiong, Ya Zhang 0002, Xue-wen Chen 0001
IEEE ACM Trans. Comput. Biol. Bioinform.2
2006 Biclustering Protein Complex Interactions with a Biclique Finding Algorithm
abstract
Biclustering has many applications in text mining, Web clickstream mining, and bioinformatics. When data entries are binary, the tightest biclusters become bicliques. We propose a flexible and highly efficient algorithm to compute bicliques. We first generalize the Motzkin-Straus formalism for computing the maximal clique from L1constraint to Lpconstraint, which enables us to provide a generalized Motzkin-Straus formalism for computing maximal-edge bicliques. By adjusting parameters, the algorithm can favor biclusters with more rows less columns, or vice verse, thus increasing the flexibility of the targeted biclusters. We then propose an algorithm to solve the generalized Motzkin-Straus optimization problem. The algorithm is provably convergent and has a computational complexity of O(/E/) where /E/ is the number of edges. Using this algorithm, we bicluster the yeast protein complex interaction network. We find that biclustering protein complexes at the protein level does not clearly reflect the functional linkage among protein complexes in many cases, while biclustering at protein domain level can reveal many underlying linkages. We show several new biologically significant results.
Chris Ding, Ya Zhang 0002, Tao Li 0001, Stephen R. Holbrook
ICDM2
2006 Motif Discovery as a Multiple-Instance Problem
abstract
Motif discovery from bio sequences, a challenging task both experimentally and computationally, has been a topic of immense study in recent years. In this paper, we formulate the motif discovery problem as a multiple-instance problem and employ a multiple-instance learning method, the MILES method, to identify motif from biological sequences. Each sequence is mapped into a feature space defined by instances in training sequences with a novel instance-bag similarity measure. We employ I-norm SVM to select important features and construct classifiers simultaneously. These high-ranked features correspond to discovered motifs. We apply this method to discover transcriptional factor binding sites in promoters, a typical motif finding problem in biology, and show that the method is at least comparable to existing methods
Ya Zhang 0002, Yixin Chen 0002, Xiang Ji 0001
ICTAI1
2006 Splice site prediction using support vector machines with a Bayes kernel
Ya Zhang 0002, Chao-Hsien Chu, Yixin Chen 0002, Hongyuan Zha, Xiang Ji 0001
Expert Syst. Appl.1
2005 Towards discovering organizational structure from email corpus
abstract
Email logs people's communication history which provides valuable information regarding the infrastructure of an organization. In this paper, a two-phase framework is introduced to attack the problem of leadership discovery in an organization based on email communication history among the employees. Two heuristic metrics are proposed for evaluating pair-wise leadership factors among a group of employees. We also address several issues in discovering the organization's structure through mining leadership graph constructed from the leadership factors. Experimental studies are carried out by applying the framework to Enron email corpus.
Yang Song 0008, Hongyuan Zha, Ya Zhang 0002
ICMLA4
2005 Size Regularized Cut for Data Clustering
abstract
We present a novel spectral clustering method that enables users to incorporate prior knowledge of the size of clusters into the clustering process. The cost function, which is named size regularized cut (SRcut), is defined as the sum of the inter-cluster similarity and a regularization term measuring the relative size of two clusters. Finding a partition of the data set to minimize SRcut is proved to be NP-complete. An approximation algorithm is proposed to solve a relaxed version of the optimization problem as an eigenvalue problem. Evaluations over different data sets demonstrate that the method is not sensitive to outliers and performs better than normalized cut.
Yixin Chen 0002, Ya Zhang 0002, Xiang Ji 0001
NIPS2
2005 Comparative mapping of sequence-based and structure-based protein domains
abstract
BACKGROUND: Protein domains have long been an ill-defined concept in biology. They are generally described as autonomous folding units with evolutionary and functional independence. Both structure-based and sequence-based domain definitions have been widely used. But whether these types of models alone can capture all essential features of domains is still an open question. METHODS: Here we provide insight on domain definitions through comparative mapping of two domain classification databases, one sequence-based (Pfam) and the other structure-based (SCOP). A mapping score is defined to indicate the significance of the mapping, and the properties of the mapping matrices are studied. RESULTS: The mapping results show a general agreement between the two databases, as well as many interesting areas of disagreement. In the cases of disagreement, the functional and evolutionary characteristics of the domains are examined to determine which domain definition is biologically more informative.
Ya Zhang 0002, John-Marc Chandonia, Chris Ding, Stephen R. Holbrook
BMC Bioinform.1
2002 Progressive display of very high resolution images using wavelets
Ya Zhang 0002, James Z. Wang 0001
AMIA1