Jun Xiao 0001

dblp:71/2308-1 · DBLP profile ↗
← Back
184ranked-venue papers
8as first author
104since 2021 · last 2026
0000-0002-6142-9914ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 112 · 8 first-author · 59 since 2021Artificial intelligence and machine learning · 88 · 1 first-author · 59 since 2021Databases, data management, data science and information retrieval · 18 · 6 since 2021Computer networks · 6 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TarPro: Targeted Protection Against Malicious Image Editing
Kaixin Shen, Ruijie Quan, Jiaxu Miao, Jun Xiao 0001
AAAI4
2026 GUI-G²: Gaussian Reward Modeling for GUI Grounding
abstract
Graphical User Interface (GUI) grounding maps natural language instructions to precise interface locations for autonomous interaction. Current reinforcement learning approaches use binary rewards that treat elements as hit-or-miss targets, creating sparse signals that ignore the continuous nature of spatial interactions. Motivated by human clicking behavior that naturally forms Gaussian distributions centered on target elements, we introduce GUI Gaussian Grounding Rewards (GUI-G2), a principled reward framework that models GUI elements as continuous Gaussian distributions across the interface plane. GUI-G2 incorporates two synergistic mechanisms: Gaussian point rewards model precise localization through exponentially decaying distributions centered on element centroids, while coverage rewards assess spatial alignment by measuring the overlap between predicted Gaussian distributions and target regions. To handle diverse element scales, we develop an adaptive variance mechanism that calibrates reward distributions based on element dimensions. This framework transforms GUI grounding from sparse binary classification to dense continuous optimization, where Gaussian distributions generate rich gradient signals that guide models toward optimal interaction positions. Extensive experiments across ScreenSpot, ScreenSpot-v2, and ScreenSpot-Pro benchmarks demonstrate that GUI-G2, substantially outperforms state-of-the-art method UI-TARS-72B, with the most significant improvement of 24.7% on ScreenSpot-Pro. Our analysis reveals that continuous modeling provides superior robustness to interface variations and enhanced generalization to unseen layouts, establishing a new paradigm for spatial reasoning in GUI interaction tasks.
Fei Tang 0005, Zhangxuan Gu, Zhengxi Lu, Shuheng Shen, Changhua Meng, Wen Wang 0009, Wenqi Zhang 0001, Yongliang Shen 0001, Weiming Lu 0001, Jun Xiao 0001, Yueting Zhuang
AAAI11
2026 MAU-GPT: Enhancing Multi-type Industrial Anomaly Understanding via Anomaly-aware and Generalist Experts Adaptation
abstract
As industrial manufacturing scales, automating fine-grained product image analysis has become critical for quality control. However, existing approaches are hindered by limited dataset coverage and poor model generalization across diverse and complex anomaly patterns. To address these challenges, we introduce MAU-Set, a comprehensive dataset for Multi-type industrial Anomaly Understanding. It spans multiple industrial domains and features a hierarchical task structure, ranging from binary classification to complex reasoning. Alongside this dataset, we establish a rigorous evaluation protocol to facilitate fair and comprehensive model assessment. Building upon this foundation, we further present MAU-GPT, a domain-adapted multimodal large model specifically designed for industrial anomaly understanding. It incorporates a novel AMoE-LoRA mechanism that unifies anomaly-aware and generalist experts adaptation, enhancing both understanding and reasoning across diverse defect classes. Extensive experiments show that MAU-GPT consistently outperforms prior state-of-the-art methods across all domains, demonstrating strong potential for scalable and automated industrial inspection.
Zhuonan Wang, Zhenxuan Fan, Siwen Tan, Yuqian Yuan, Haoyuan Li 0002, Hao Jiang 0014, Wenqiao Zhang, Feifei Shao, Jun Xiao 0001
AAAI11
2026 UI-Copilot: Advancing Long-Horizon GUI Automation via Tool-Integrated Policy Optimization
abstract
Zhengxi Lu, Fei Tang, Guangyi Liu, Jin Ma, Kaitao Song, Xu Tan, Wenqi Zhang, Weiming Lu, Jun Xiao, Yueting Zhuang, Yongliang Shen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhengxi Lu, Fei Tang 0005, Kaitao Song, Xu Tan 0003, Wenqi Zhang 0001, Weiming Lu 0001, Jun Xiao 0001, Yueting Zhuang, Yongliang Shen 0001
ACL (1)9
2026 Experience-driven Multi-turn Reinforcement Learning for GUI Agents
abstract
Zhengxi Lu, Jiabo Ye, Fei Tang, Yongliang Shen, Haiyang Xu, Ziwei Zheng, Weiming Lu, Ming Yan, Fei Huang, Jun Xiao, Yueting Zhuang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhengxi Lu, Jiabo Ye, Fei Tang 0005, Yongliang Shen 0001, Haiyang Xu 0001, Ziwei Zheng, Weiming Lu 0001, Ming Yan 0008, Fei Huang 0002, Jun Xiao 0001, Yueting Zhuang
ACL (1)10
2026 CoVerRL: Breaking the Consensus Trap in Label-Free Reasoning via Generator-Verifier Co-Evolution
abstract
Teng Pan, Yuchen Yan, Zixuan Wang, Ruiqing Zhang, Guiyang Hou, Wenqi Zhang, Weiming Lu, Jun Xiao, Yongliang Shen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Teng Pan, Ruiqing Zhang, Guiyang Hou, Wenqi Zhang 0001, Weiming Lu 0001, Jun Xiao 0001, Yongliang Shen 0001
ACL (1)8
2026 PILOT: Planning via Internalized Latent Optimization Trajectories for Large Language Models
abstract
Haoyu Zheng, Yun Zhu, Yuqian Yuan, Bo Yuan, Wenqiao Zhang, Siliang Tang, Jun Xiao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yun Zhu 0007, Yuqian Yuan, Wenqiao Zhang, Siliang Tang, Jun Xiao 0001
ACL (1)7
2026 Audio-Guided Video Scene Editing
Kaixin Shen, Ruijie Quan, Linchao Zhu, Jun Xiao 0001, Yi Yang 0001
Int. J. Comput. Vis.5
2026 Momentor++: Advancing Video Large Language Models With Fine-Grained Long Video Reasoning
abstract
Large Language Models (LLMs) exhibit remarkable proficiency in understanding and managing text-based tasks.Many works try to transfer these capabilities to the video domain, which are referred to as Video-LLMs. However, current Video-LLMs can only grasp the coarse-grained semantics and are unable to efficiently handle tasks involving the comprehension or localization of specific video segments. To address these challenges, we propose Momentor, a Video-LLM designed to perform fine-grained temporal understanding tasks. To facilitate the training of Momentor, we develop an automatic data generation engine to build Moment-10M, a large-scale video instruction dataset with segment-level instruction data. Building upon the foundation of the previously published Momentor and the Moment-10M dataset, we further extend this work by introducing a Spatio-Temporal Token Consolidation (STTC) method, which can merge redundant visual tokens spatio-temporally in a parameter-free manner, thereby significantly promoting computational efficiency while preserving fine-grained visual details. We integrate STTC with Momentor to develop Momentor++ and validate its performance on various benchmarks. Momentor demonstrates robust capabilities in fine-grained temporal understanding and localization. Further, Momentor++ excels in efficiently processing and analyzing extended videos with complex events, showcasing marked advancements in handling extensive temporal contexts.
Juncheng Li 0006, Minghe Gao, Xiangnan He 0001, Siliang Tang, Wei-Shi Zheng 0001, Jun Xiao 0001, Meng Wang 0001, Tat-Seng Chua, Yueting Zhuang
IEEE Trans. Pattern Anal. Mach. Intell.6
2026 Structure-Induced Gradient Regulation for Generalizable Vision-Language Models
abstract
Prompt tuning, a recently emerging paradigm, adapts vision-language pre-trained models to new tasks efficiently by learning "soft prompts" for frozen models. However, in few-shot scenarios, its effectiveness is limited by sensitivity to the initialization and the time-consuming search for optimal initialization, hindering rapid adaptation. Additionally, prompt tuning risks reducing the models' generalizability due to overfitting on scarce training samples. To overcome these challenges, we introduce a novel Gradient-RegulAted Meta-prompt learning (GRAM) framework that jointly meta-learns an efficient soft prompt initialization for better adaptation and a lightweight gradient regulating function for strong cross-domain generalizability in a meta-learning paradigm using only the weakly labeled image-text pre-training data. This is achieved through a Cross-Modal Hierarchical Clustering algorithm that organizes extensive image-text data into a structured hierarchy, facilitating robust meta-learning across diverse domains. Rather than designing a specific prompt tuning method, our GRAM can be easily incorporated into various prompt tuning methods in a model-agnostic way and bring about consistent improvement for them. Further, we consider a more practical but challenging setting: test-time prompt tuning with only unlabeled test samples and propose an improved structure-induced gradient regulating function to leverage the structured semantics of the meta-learning data for zero-shot generalization. This novel approach exploits the hierarchically clustered meta-learning data to model relationships between test-time data and meta-learning prototypes, facilitating the transfer of invariant knowledge without explicit annotations. Meanwhile, we introduce a structure complexity-informed strategy for adaptively constructing meta-training tasks and generating prototypes, which fully considers the diverse semantics within hierarchical clusters of different complexities. Comprehensive experiments demonstrate the state-of-the-art few- and zero-shot generalizability of our method.
Juncheng Li 0006, Minghe Gao, Siliang Tang, Longhui Wei, Jun Xiao 0001, Fei Wu 0001, Richang Hong, Meng Wang 0001, Qi Tian 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 Improving Model Fusion by Training-Time Neuron Alignment With Fixed Neuron Anchors
abstract
Model fusion aims to integrate several deep neural network (DNN) models' knowledge into one by fusing parameters, and it has promising applications, such as improving the generalization of foundation models and parameter averaging in federated learning. However, models under different settings (data, hyperparameter, etc.) have diverse neuron permutations; in other words, from the perspective of loss landscape, they reside in different loss basins, thus hindering model fusion performances. To alleviate this issue, previous studies highlighted the role of permutation invariance and have developed methods to find correct network permutations for neuron alignment after training. Orthogonal to previous attempts, this paper studies training-time neuron alignment, improving model fusion without the need for post-matching. Training-time alignment is cheaper than post-alignment and is applicable in various model fusion scenarios. Starting from fundamental hypotheses and theorems, a simple yet lossless algorithm called TNA-PFN is introduced. TNA-PFN utilizes partially fixed neuron weights as anchors to reduce the potential of training-time permutations, and it is empirically validated in reducing the barriers of linear mode connectivity and multi-model fusion. It is also validated that TNA-PFN can improve the fusion of pretrained models under the setting of model soup (vision transformers) and ColD fusion (pretrained language models). Based on TNA-PFN, two federated learning methods, FedPFN and FedPNU, are proposed, showing the prospects of training-time neuron alignment. FedPFN and FedPNU reach state-of-the-art performances in federated learning under heterogeneous settings and can be compatible with the server-side algorithm.
Zexi Li 0001, Zhiqi Li 0004, Tao Shen 0002, Jun Xiao 0001, Yike Guo, Tao Lin 0004, Chao Wu 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 Generalized Visual Relation Detection With Diffusion Models
abstract
Visual relation detection (VRD) aims to identify relationships (or interactions) between object pairs in an image. Although recent VRD models have achieved impressive performance, they are all restricted to pre-defined relation categories, while failing to consider thesemantic ambiguitycharacteristic of visual relations. Unlike objects, the appearance of visual relations is always subtle and can be described by multiple predicate words from different perspectives, e.g., “ride” can be depicted as “race” and “sit on”, from the sports and spatial position views, respectively. To this end, we propose to model visual relations as continuous embeddings, and design diffusion models to achieve generalized VRD in a conditional generative manner, termed Diff-VRD. We model the diffusion process in a latent space and generate all possible relations in the image as an embedding sequence. During the generation, the visual and text embeddings of subject-object pairs serve as conditional signals and are injected via cross-attention. After the generation, we design a subsequent matching stage to assign the relation words to subject-object pairs by considering their semantic similarities. Benefiting from the diffusion-based generative process, our Diff-VRD is able to generate visual relations beyond the pre-defined category labels of datasets. To properly evaluate this generalized VRD task, we introduce two evaluation metrics, i.e., text-to-image retrieval and SPICE PR Curve inspired by image captioning. Extensive experiments in both human-object interaction (HOI) detection and scene graph generation (SGG) benchmarks attest to the superiority and effectiveness of Diff-VRD.
Kaifeng Gao, Hanwang Zhang, Jun Xiao 0001, Yueting Zhuang, Qianru Sun
IEEE Trans. Circuits Syst. Video Technol.4
2026 Physically Plausible Human-Object Rendering From Sparse Views via 3D Gaussian Splatting
abstract
Rendering realistic human-object interactions (HOIs) from sparse-view inputs is a challenging yet crucial task for various real-world applications. Existing methods often struggle to simultaneously achieve high rendering quality, physical plausibility, and computational efficiency. To address these limitations, we propose HOGS (Human-Object Rendering via 3D Gaussian Splatting), a novel framework for efficient HOI rendering with physically plausible geometric constraints from sparse views. HOGS represents both humans and objects as dynamic 3D Gaussians. Central to HOGS is a novel optimization process that operates directly on these Gaussians to enforce geometric consistency (i.e., preventing inter-penetration or floating contacts) to achieve physical plausibility. To support this core optimization under sparse-view ambiguity, our framework incorporates two pre-trained modules: an optimization-guided Human Pose Refiner for robust estimation under sparse-view occlusions, and a Human-Object Contact Predictor that efficiently identifies interaction regions to guide our novel contact and separation losses. Extensive experiments on both human-object and hand-object interaction datasets demonstrate that HOGS achieves state-of-the-art rendering quality and maintains high computational efficiency.
Jun Xiao 0001, Yi Yang 0001, Yueting Zhuang, Long Chen 0016
IEEE Trans. Image Process.2
2026 Counterfactual Co-Occurring Learning for Bias Mitigation in Weakly-Supervised Object Localization
abstract
Contemporary weakly-supervised object localization (WSOL) methods have primarily focused on addressing the challenge of localizing the most discriminative region while largely overlooking the relatively less explored issue of biased activation—incorrectly spotlighting co-occurring background with the foreground feature. In this paper, we conduct a thorough causal analysis to investigate the origins of biased activation. Based on our analysis, we attribute this phenomenon to the presence of co-occurring background confounders. Building upon this profound insight, we introduce a pioneering paradigm known as Counterfactual Co-occurring Learning (CCL), meticulously engendering counterfactual representations by adeptly disentangling the foreground from the co-occurring background elements. Furthermore, we propose an innovative network architecture known as Counterfactual-CAM. This architecture seamlessly incorporates a perturbation mechanism for counterfactual representations into the vanilla CAM-based model. By training the WSOL model with these perturbed representations, we guide the model to prioritize the consistent foreground content while concurrently reducing the influence of distracting co-occurring backgrounds. To the best of our knowledge, this study represents the initial exploration of this research direction. Our extensive experiments conducted across multiple benchmarks validate the effectiveness of the proposed Counterfactual-CAM in mitigating biased activation.
Feifei Shao, Yawei Luo, Lei Chen 0082, Ping Liu 0004, Wei Yang 0034, Yi Yang 0001, Jun Xiao 0001
IEEE Trans. Multim.7
2026 FreeTuner: Any Subject in Any Style with Training-free Diffusion
abstract
With the advance of diffusion models, various personalized image generation methods have been proposed. However, almost all existing work only focuses on either subject-driven or style-driven personalization. Meanwhile, state-of-the-art methods face several challenges in realizing compositional personalization , that is, composing different subject and style concepts, such as concept disentanglement, unified reconstruction paradigm, and insufficient training data. To address these issues, we introduce FreeTuner , a flexible and training-free method for compositional personalization that can generate any user-provided subject in any user-provided style . Our approach employs a disentanglement strategy that separates the generation process into two stages to effectively mitigate concept entanglement. FreeTuner leverages the intermediate features within the diffusion model for subject concept representation and introduces style guidance to align the synthesized images with the style concept, ensuring the preservation of both the subject’s structure and the style’s aesthetic features. Extensive experiments have demonstrated the generation ability of FreeTuner across various personalization settings.
Youcan Xu, Zhen Wang 0004, Jun Xiao 0001, Long Chen 0016
ACM Trans. Multim. Comput. Commun. Appl.4
2025 Towards Better Alignment: Training Diffusion Models with Reinforcement Learning Against Sparse Rewards
abstract
Diffusion models have achieved remarkable success in text-to-image generation. However, their practical applications are hindered by the misalignment between generated images and corresponding text prompts. To tackle this issue, reinforcement learning (RL) has been considered for diffusion model fine-tuning. Yet, RL’s effectiveness is limited by the challenge of sparse reward, where feedback is only available at the end of the generation process. This makes it difficult to identify which actions during the de-noising process contribute positively to the final generated image, potentially leading to ineffective or unnecessary de-noising policies. To this end, this paper presents a novel RL-based framework that addresses the sparse reward problem when training diffusion models. Our framework, named B2-DiffuRL, employs two strategies: Backward progressive training and Branch-based sampling. For one thing, backward progressive training focuses initially on the final timesteps of denoising process and gradually extends the training interval to earlier timesteps, easing the learning difficulty from sparse rewards. For another, we perform branch-based sampling for each training interval. By comparing the samples within the same branch, we can identify how much the policies of the current training interval contribute to the final image, which helps to learn effective policies instead of unnecessary ones. B2-DiffuRL is compatible with existing optimization algorithms. Extensive experiments demonstrate the effectiveness of B2-DiffuRL in improving prompt-image alignment and maintaining diversity in generated images. The code for this work is available1.
Zijing Hu, Fengda Zhang, Long Chen 0016, Kun Kuang 0001, Jiahui Li 0003, Kaifeng Gao, Jun Xiao 0001, Xin Wang 0019, Wenwu Zhu 0001
CVPR7
2025 MICAS: Multi-grained In-Context Adaptive Sampling for 3D Point Cloud Processing
abstract
Point cloud processing (PCP) encompasses tasks like reconstruction, denoising, registration, and segmentation, each often requiring specialized models to address unique task characteristics. While in-context learning (ICL) has shown promise across tasks by using a single model with task-specific demonstration prompts, its application to PCP reveals significant limitations. We identify inter-task and intra-task sensitivity issues in current ICL methods for PCP, which we attribute to inflexible sampling strategies lacking context adaptation at the point and prompt levels. To address these challenges, we propose MICAS, an advanced ICL framework featuring a multi-grained adaptive sampling mechanism tailored for PCP. MICAS introduces two core components: task-adaptive point sampling, which leverages inter-task cues for point-level sampling, and query-specific prompt sampling, which selects optimal prompts per query to mitigate intra-task sensitivity. To our knowledge, this is the first approach to introduce adaptive sampling tailored to the unique requirements of point clouds within an ICL framework. Extensive experiments show that MICAS not only efficiently handles various PCP tasks but also significantly outperforms existing methods. Notably, it achieves a remarkable 4.1% improvement in the part segmentation task and delivers consistent gains across various PCP applications.
Feifei Shao, Ping Liu 0004, Yawei Luo, Jun Xiao 0001
CVPR6
2025 TAGA: Self-supervised Learning for Template-free Animatable Gaussian Articulated Model
abstract
Decoupling from customized parametric templates represents a crucial step toward the creation of fully flexible, animatable articulated models. While existing template-free methods can achieve high-fidelity reconstruction in observed views, they struggle to recover plausible canonical models, resulting in suboptimal animation quality. This limitation stems from overlooking the fundamental ambiguity in canonical reconstruction, where multiple canonical models could explain the same observed views. By revealing the entanglement between the canonical ambiguity and incorrect skinning, we present a self-supervised framework that learns both plausible skinning and accurate canonical geometry using only sparse pose data. Our method, TAGA, uses explicit 3D Gaussians as skinning carriers and characterizes the ambiguity as "Ambiguous Gaussians" with incorrect skinning weights. TAGA then corrects ambiguous Gaussians in the observation space using anomaly detection. With the corrected ones, we enforce cycle consistency constraints on both geometry and skinning to refine the corresponding Gaussians in the canonical space through a new backward method. Compared to existing state-of-the-art template-free methods, TAGA delivers superior visual fidelity for novel views and poses, while significantly improving training and rendering speeds. Experiments on challenging datasets with limited pose variations further demonstrate the robustness and generality of TAGA.
Zhichao Zhai, Guikun Chen, Wenguan Wang, Jun Xiao 0001
CVPR5
2025 Decoding Correlation-Induced Misalignment in the Stable Diffusion Workflow for Text-to-Image Generation
Yunze Tong, Fengda Zhang, Didi Zhu, Jun Xiao 0001, Kun Kuang 0001
ICCV4
2025 Event-Customized Image Generation
abstract
Customized Image Generation, generating customized images with user-specified concepts, has raised significant attention due to its creativity and novelty. With impressive progress achieved in subject customization, some pioneer works further explored the customization of action and interaction beyond entity (i.e., human, animal, and object) appearance. However, these approaches only focus on basic actions and interactions between two entities, and their effects are limited by insufficient ”exactly same” reference images. To extend customized image generation to more complex scenes for general real-world applications, we propose a new task: event-customized image generation. Given a single reference image, we define the ”event” as all specific actions, poses, relations, or interactions between different entities in the scene. This task aims at accurately capturing the complex event and generating customized images with various target entities. To solve this task, we proposed a novel training-free event customization method: FreeEvent. Specifically, FreeEvent introduces two extra paths alongside the general diffusion denoising process: 1) Entity switching path: it applies cross-attention guidance and regulation for target entity generation. 2) Event transferring path: it injects the spatial feature and self-attention maps from the reference image to the target image for event generation. To further facilitate this new task, we collected two evaluation benchmarks: SWiG-Event and Real-Event. Extensive experiments and ablations have demonstrated the effectiveness of FreeEvent.
Zhen Wang 0004, Yilei Jiang, Jun Xiao 0001, Long Chen 0016
ICML4
2025 Ca2-VDM: Efficient Autoregressive Video Diffusion Model with Causal Generation and Cache Sharing
abstract
With the advance of diffusion models, today's video generation has achieved impressive quality. To extend the generation length and facilitate real-world applications, a majority of video diffusion models (VDMs) generate videos in an autoregressive manner, i.e., generating subsequent clips conditioned on the last frame(s) of the previous clip. However, existing autoregressive VDMs are highly inefficient and redundant: The model must re-compute all the conditional frames that are overlapped between adjacent clips. This issue is exacerbated when the conditional frames are extended autoregressively to provide the model with long-term context. In such cases, the computational demands increase significantly (i.e., with a quadratic complexity w.r.t. the autoregression step). In this paper, we propose **Ca2-VDM**, an efficient autoregressive VDM with **Ca**usal generation and **Ca**che sharing. For **causal generation**, it introduces unidirectional feature computation, which ensures that the cache of conditional frames can be precomputed in previous autoregression steps and reused in every subsequent step, eliminating redundant computations. For **cache sharing**, it shares the cache across all denoising steps to avoid the huge cache storage cost. Extensive experiments demonstrated that our Ca2-VDM achieves state-of-the-art quantitative and qualitative video generation results and significantly improves the generation speed. Code is available: https://github.com/Dawn-LX/CausalCache-VDM
Kaifeng Gao, Jiaxin Shi, Hanwang Zhang, Chunping Wang 0001, Jun Xiao 0001, Long Chen 0016
ICML5
2025 HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation
abstract
We present **HealthGPT**, a powerful Medical Large Vision-Language Model (Med-LVLM) that integrates medical visual comprehension and generation capabilities within a unified autoregressive paradigm. Our bootstrapping philosophy is to progressively adapt heterogeneous comprehension and generation knowledge to pre-trained Large Language Models (LLMs). This is achieved through a novel heterogeneous low-rank adaptation **(H-LoRA)** technique, which is complemented by a tailored hierarchical visual perception **(HVP)** approach and a three-stage learning strategy **(TLS)**. To effectively learn the HealthGPT, we devise a comprehensive medical domain-specific comprehension and generation dataset called **VL-Health**. Experimental results demonstrate exceptional performance and scalability of HealthGPT in medical visual unified tasks. Our project can be accessed at https://github.com/DCDmllm/HealthGPT.
Tianwei Lin 0001, Wenqiao Zhang, Sijing Li, Yuqian Yuan, Binhe Yu, Haoyuan Li 0002, Wanggui He, Hao Jiang 0014, Mengze Li 0001, Siliang Tang, Jun Xiao 0001, Yueting Zhuang, Beng Chin Ooi
ICML12
2025 Latent Score-Based Reweighting for Robust Classification on Imbalanced Tabular Data
abstract
Machine learning models often perform well on tabular data by optimizing average prediction accuracy. However, they may underperform on specific subsets due to inherent biases and spurious correlations in the training data, such as associations with non-causal features like demographic information. These biases lead to critical robustness issues as models may inherit or amplify them, resulting in poor performance where such misleading correlations do not hold. Existing mitigation methods have significant limitations: some require prior group labels, which are often unavailable, while others focus solely on the conditional distribution $P(Y|X)$, upweighting misclassified samples without effectively balancing the overall data distribution $P(X)$. To address these shortcomings, we propose a latent score-based reweighting framework. It leverages score-based models to capture the joint data distribution $P(X, Y)$ without relying on additional prior information. By estimating sample density through the similarity of score vectors with neighboring data points, our method identifies underrepresented regions and upweights samples accordingly. This approach directly tackles inherent data imbalances, enhancing robustness by ensuring a more uniform dataset representation. Experiments on various tabular datasets under distribution shifts demonstrate that our method effectively improves performance on imbalanced data.
Yunze Tong, Fengda Zhang, Kaifeng Gao, Pengfei Lyu, Jun Xiao 0001, Kun Kuang 0001
ICML7
2025 Compositional Zero-shot Learning via Progressive Language-based Observations
abstract
Compositional zero-shot learning aims to recognize unseen stateobject compositions by leveraging known primitives (state and object) during training. However, effectively modeling interactions between primitives and generalizing knowledge to novel compositions remains a perennial challenge. There are two crucial factors: large object-conditioned and state-conditioned variance, i.e., the appearance of states (or objects) can vary significantly when combined with different objects (or states). For instance, the state "old" can signify vintage design for a "car" or advanced age for a "cat". In this paper, we argue that these variances can be mitigated by predicting composition categories based on salient observation cues. Therefore, we propose Progressive Language-based Observations (PLO), which can automatically determine the order of observation cues. These "observation cues" comprise a series of primitive concepts or graduated descriptions that allow the model to understand image content in a step-by-step manner. Specifically, PLO adopts pre-trained vision-language models (VLMs) to empower the model with observation capabilities.We further devise two variants: a twostep method (PLO-VLM) with a pre-observing classifier dynamically selecting the order of primitive concept-based cues, and a multistep approach (PLO-LLM) using large language models (LLMs) to craft graduated description-based cues. Extensive tests on three datasets show PLO's effectiveness in compositional recognition.
Lin Li 0065, Guikun Chen, Zhen Wang 0004, Jun Xiao 0001, Long Chen 0016
ACM Multimedia4
2025 EyecareGPT: Boosting Comprehensive Ophthalmology Understanding with Tailored Dataset, Benchmark and Model
Sijing Li, Tianwei Lin 0001, Lingshuai Lin, Wenqiao Zhang, Xiaoda Yang, Juncheng Li 0006, Jun Xiao 0001, Yueting Zhuang, Beng Chin Ooi
ACM Multimedia10
2025 Robust Modality-Incomplete Anomaly Detection: A Modality-Instructive Framework with Benchmark
abstract
Multimodal Industrial Anomaly Detection (MIAD)-fusing 3D point clouds and 2D RGB for product defect detection-is critical to quality inspection. However, existing MIAD methods assume all modalities are available and paired, overlooking real-scenario modality-missing and risking overfitting to incomplete data. To address these, we conduct the first comprehensive study on Modality-Incomplete Industrial Anomaly Detection (MIIAD) and establish MIIAD Bench , a benchmark covering diverse missing settings. Meanwhile, we propose RADAR, a robust two-stage Robust modAlity-instructive fusing & Detecting frAmewoRk. RADAR integrates i) a Modality-Incomplete Instruction mechanism-guiding the multimodal Transformer to focus more on available modal info, and ii) a Double-Pseudo Hybrid Module to highlight unique modality combinations and reduce overfitting. Our results show RADAR outperforms prior methods markedly on MIIAD Bench.
Bingchen Miao, Wenqiao Zhang, Juncheng Li 0006, Wangyu Wu, Siliang Tang, Zhaocheng Li, Jun Xiao 0001, Yueting Zhuang
ACM Multimedia8
2025 Zero-shot Compositional Action Recognition with Neural Logic Constraints
abstract
Zero-shot compositional action recognition (ZS-CAR) aims to identify unseen verb-object compositions in the videos by exploiting the learned knowledge of verb and object primitives during training. Despite compositional learning's progress in ZS-CAR, two critical challenges persist: 1) Missing compositional structure constraint, leading to spurious correlations between primitives; 2) Neglecting semantic hierarchy constraint, leading to semantic ambiguity and impairing the training process. In this paper, we argue that human-like symbolic reasoning offers a principled solution to these challenges by explicitly modeling compositional and hierarchical structured abstraction. To this end, we propose a logic-driven ZS-CAR framework LogicCAR that integrates dual symbolic constraints: Explicit Compositional Logic and Hierarchical Primitive Logic. Specifically, the former models the restrictions within the compositions, enhancing the compositional reasoning ability of our model. The latter investigates the semantical dependencies among different primitives, empowering the models with fine-to-coarse reasoning capacity. By formalizing these constraints in first-order logic and embedding them into neural network architectures, LogicCAR systematically bridges the gap between symbolic abstraction and existing models. Extensive experiments on the Sth-com dataset demonstrate that our LogicCAR outperforms existing baseline methods, proving the effectiveness of our logic-driven constraints.
Gefan Ye, Lin Li 0065, Jun Xiao 0001, Long Chen 0016
ACM Multimedia4
2025 Counterfactual Evolution of Multimodal Datasets via Visual Programming
abstract
The rapid development of Multimodal Large Language Models (MLLMs) poses increasing demands on the diversity and complexity of multimodal datasets. Yet manual annotation pipelines can no longer keep pace. Existing augmentation methods often follow fixed rules and lack verifiable control over sample diversity and reasoning complexity. To address this, we introduce Scalable COunterfactual Program Evolution (SCOPE), a framework that uses symbolic Visual Programming to guide program evolution via counterfactual reasoning. SCOPE performs the three steps of counterfactual inference: (1) Abduction, by generating verifiable programs to model reasoning associations; (2) Action, by intervening on program structure along three axes—reasoning path, visual context, and cross-instance composition; and (3) Prediction, by categorizing evolved instances by difficulty, structure, and input multiplicity. Based on this process, we build SCOPE-Train and SCOPE-Test, evolving benchmarks with expert validation. To support training, we propose MAP, a curriculum learning strategy that aligns model capacity with sample difficulty. Experiments show that SCOPE improves reasoning performance, exposes model blind spots, and enhances visual dialog capabilities.
Minghe Gao, Zhongqi Yue, Wei Ji 0008, Siliang Tang, Jun Xiao 0001, Tat-Seng Chua, Yueting Zhuang, Juncheng Li 0006
NeurIPS7
2025 Janus-Pro-R1: Advancing Collaborative Visual Comprehension and Generation via Reinforcement Learning
abstract
Recent endeavors in Multimodal Large Language Models (MLLMs) aim to unify visual comprehension and generation. However, these two capabilities remain largely independent, as if they are two separate functions encapsulated within the same model. Consequently, visual comprehension does not enhance visual generation, and the reasoning mechanisms of LLMs have not been fully integrated to revolutionize image generation. In this paper, we propose to enable the collaborative co-evolution of visual comprehension and generation, advancing image generation into an iterative introspective process. We introduce a two-stage training approach: supervised fine-tuning teaches the MLLM with the foundational ability to generate genuine CoT for visual generation, while reinforcement learning activates its full potential via an exploration-exploitation trade-off. Ultimately, we unlock the Aha moment in visual generation, advancing MLLMs from text-to-image tasks to unified image generation. Extensive experiments demonstrate that our model not only excels in text-to-image generation and image editing, but also functions as a superior image semantic evaluator with enhanced visual comprehension capabilities. Project Page: \url{https://janus-pro-r1.github.io}.
Kaihang Pan, Wendong Bu, Juncheng Li 0006, Yingting Wang, Siliang Tang, Jun Xiao 0001, Fei Wu 0001, Yueting Zhuang
NeurIPS9
2025 Mind the Gap: Bridging Thought Leap for Improved Chain-of-Thought Tuning
abstract
Large language models (LLMs) have achieved remarkable progress on mathematical tasks through Chain-of-Thought (CoT) reasoning. However, existing mathematical CoT datasets often suffer from **Thought Leaps** due to experts omitting intermediate steps, which negatively impacts model learning and generalization. We propose the CoT Thought Leap Bridge Task, which aims to automatically detect leaps and generate missing intermediate reasoning steps to restore the completeness and coherence of CoT. To facilitate this, we constructed a specialized training dataset called **ScaleQM+**, based on the structured ScaleQuestMath dataset, and trained **CoT-Bridge** to bridge thought leaps. Through comprehensive experiments on mathematical reasoning benchmarks, we demonstrate that models fine-tuned on bridged datasets consistently outperform those trained on original datasets, with improvements of up to +5.87\% on NuminaMath. Our approach effectively enhances distilled data (+3.02\%) and provides better starting points for reinforcement learning (+3.1\%), functioning as a plug-and-play module compatible with existing optimization techniques. Furthermore, CoT-Bridge demonstrates improved generalization to out-of-domain logical reasoning tasks, confirming that enhancing reasoning completeness yields broadly applicable benefits.
Haolei Xu, Yongliang Shen 0001, Wenqi Zhang 0001, Guiyang Hou, Shengpei Jiang, Kaitao Song, Weiming Lu 0001, Jun Xiao 0001, Yueting Zhuang
NeurIPS9
2025 EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World?
abstract
The emergence of multimodal large language models (MLLMs) has driven breakthroughs in egocentric vision applications. These applications necessitate persistent, context-aware understanding of objects, as users interact with tools in dynamic and cluttered environments. However, existing embodied benchmarks primarily focus on static scene exploration, emphasizing object's appearance and spatial attributes while neglecting the assessment of dynamic changes arising from users' interactions.capabilities in object-level spatiotemporal reasoning required for real-world interactions.To address this gap, we introduce EOC-Bench, an innovative benchmark designed to systematically evaluate object-centric embodied cognition in dynamic egocentric scenarios.Specially, EOC-Bench features 3,277 meticulously annotated QA pairs categorized into three temporal categories: Past, Present, and Future, covering 11 fine-grained evaluation dimensions and 3 visual object referencing types.To ensure thorough assessment, we develop a mixed-format human-in-the-loop annotation frameworkBased on EOC-Bench, we conduct comprehensive evaluations of various proprietary, open-source, and object-level MLLMs. EOC-Bench serves as a crucial tool for advancing the embodied object cognitive capabilities of MLLMs, establishing a robust foundation for developing reliable core models for embodied systems.
Yuqian Yuan, Ronghao Dang, Wentong Li 0001, Xin Li 0056, Deli Zhao, Fan Wang 0019, Wenqiao Zhang, Jun Xiao 0001, Yueting Zhuang
NeurIPS10
2025 Let LRMs Break Free from Overthinking via Self-Braking Tuning
abstract
Large reasoning models (LRMs), such as OpenAI o1 and DeepSeek-R1, have significantly enhanced their reasoning capabilities by generating longer chains of thought, demonstrating outstanding performance across a variety of tasks. However, this performance gain comes at the cost of a substantial increase in redundant reasoning during the generation process, leading to high computational overhead and exacerbating the issue of overthinking. Although numerous existing approaches aim to address the problem of overthinking, they often rely on external interventions. In this paper, we propose a novel framework, **Self-Braking Tuning**(SBT), which tackles overthinking from the perspective of allowing the model to regulate its own reasoning process, thus eliminating the reliance on external control mechanisms. We construct a set of overthinking identification metrics based on standard answers and design a systematic method to detect redundant reasoning. This method accurately identifies unnecessary steps within the reasoning trajectory and generates training signals for learning self-regulation behaviors. Building on this foundation, we develop a complete strategy for constructing data with adaptive reasoning lengths and introduce an innovative braking prompt mechanism that enables the model to naturally learn when to terminate reasoning at an appropriate point. Experiments across mathematical benchmarks (AIME, AMC, MATH500, GSM8K) demonstrate that our method reduces token consumption by up to 60\% while maintaining comparable accuracy to unconstrained models.
Yongliang Shen 0001, Haolei Xu, Wenqi Zhang 0001, Kaitao Song, Jian Shao 0001, Weiming Lu 0001, Jun Xiao 0001, Yueting Zhuang
NeurIPS9
2025 From Easy to Hard: Learning Curricular Shape-Aware Features for Robust Panoptic Scene Graph Generation
Hanrong Shi, Lin Li 0065, Jun Xiao 0001, Yueting Zhuang, Long Chen 0016
Int. J. Comput. Vis.3
2025 Learning Combinatorial Prompts for Universal Controllable Image Captioning
abstract
Abstract Controllable Image Captioning (CIC)—generating natural language descriptions about images under the guidance of given control signals—is one of the most promising directions toward next-generation captioning systems. Till now, various kinds of control signals for CIC have been proposed, ranging from content-related control to structure-related control. However, due to the format and target gaps of different control signals, all existing CIC works (or architectures) only focus on one certain control signal, and overlook the human-like combinatorial ability. By “combinatorial", we mean that our humans can easily meet multiple needs (or constraints) simultaneously when generating descriptions. To this end, we propose a novel prompt-based framework for CIC by learning Com binatorial Pro mpts, dubbed as ComPro . Specifically, we directly utilize a pretrained language model GPT-2 Radford et al. (OpenAI blog 1:9, 2019) as our language model, which can help to bridge the gap between different signal-specific CIC architectures. Then, we reformulate the CIC as a prompt-guide sentence generation problem, and propose a new lightweight prompt generation network to generate the combinatorial prompts for different kinds of control signals. For different control signals, we further design a new mask attention mechanism to realize the prompt-based CIC. Due to its simplicity, our ComPro can be further extended to more kinds of combined control signals by concatenating these prompts. Extensive experiments on two prevalent CIC benchmarks have verified the effectiveness and efficiency of our ComPro on both single and combined control signals.
Zhen Wang 0004, Jun Xiao 0001, Yueting Zhuang, Fei Gao 0014, Jian Shao 0001, Long Chen 0016
Int. J. Comput. Vis.2
2025 An adaptive outlier correction quantization method for vision Transformers
abstract
Transformers have demonstrated considerable success across various domains but are constrained by their significant computational and memory requirements. This poses challenges for deployment on resource-constrained devices. Quantization, as an effective model compression method, can significantly reduce the operational time of Transformers on edge devices. Notably, Transformers display more substantial outliers than convolutional neural networks, leading to uneven feature distribution among different channels and tokens. To address this issue, we propose an adaptive outlier correction quantization (AOCQ) method for Transformers, which significantly alleviates the adverse effects of these outliers. AOCQ adjusts the notable discrepancies in channels and tokens across three levels: operator level, framework level, and loss level. We introduce a new operator that equivalently balances the activations across different channels and insert an extra stage to optimize the activation quantization step on the framework level. Additionally, we transfer the imbalanced activations across tokens and channels to the optimization of model weights on the loss level. Based on the theoretical study, our method can reduce the quantization error. The effectiveness of the proposed method is verified on various benchmark models and tasks. Surprisingly, DeiT-Base with 8-bit post-training quantization (PTQ) can achieve 81.57% accuracy with a 0.28 percentage point drop while enjoying 4× faster runtime. Furthermore, the weights of Swin and DeiT on several tasks, including classification and object detection, can be post-quantized to ultra-low 4 bits, with a minimal accuracy loss of 2%, while requiring nearly 8× less memory.
Zheyang Li, Chaoxiang Lan, Kai Zhang 0055, Wenming Tan, Ye Ren, Jun Xiao 0001
Frontiers Inf. Technol. Electron. Eng.6
2025 Knowledge Integration for Grounded Situation Recognition
Jiaming Lei, Sijing Wu, Lin Li 0065, Lei Chen 0082, Jun Xiao 0001, Yi Yang 0001, Long Chen 0016
Pattern Recognit.5
2025 Robust Global Localization for Urban Autonomous Vehicles via 3D Geometric-Enhanced Visual Place Recognition
abstract
Accurate and robust long-term global localization is a critical challenge for autonomous vehicles operating in complex urban transportation systems, where GPS signals are often unreliable and visual/inertial odometry suffers from inevitable error accumulation. Visual Place Recognition (VPR) offers a crucial solution by detecting loop closures to mitigate trajectory drift, but its performance severely degrades under complex urban traffic scenarios, such as drastic changes in viewpoint, illumination, and weather. To address these limitations, we propose the 3D Geometric feature Enhanced VPR (GE-VPR), a novel framework that improves the robustness of vehicular global localization. As a purely vision-based system, GE-VPR reconstruct 3D point clouds through dense simultaneous localization and mapping. A 3D geometric feature extraction network is designed to obtain the stable structural features of the point clouds, and a 2D-3D hybrid network is then developed to further augment these 3D features with 2D semantics. Additionally, a descriptor refinement strategy is proposed to fine-tune the raw 2D descriptors by aggregating the most relevant hybrid structural features, thus effectively fusing rich 2D appearance, color, and semantic information with stable 3D geometry. Extensive experiments on challenging urban autonomous driving datasets demonstrate that GE-VPR significantly improves vehicle localization accuracy and robustness. The overall recognition recall is increased by more than 5%, and the positioning accuracy is significantly improved in practical scenarios, demonstrating its potential as an effective solution to improve the safety and reliability of vehicular localization in autonomous navigation systems.
Junpeng Shang, Yue Liu 0048, Jun Xiao 0001, Dongfang Ma
IEEE Trans. Intell. Transp. Syst.4
2025 ENCODE: Breaking the Trade-Off Between Performance and Efficiency in Long-Term User Behavior Modeling
abstract
Long-term user behavior sequences are a goldmine for businesses to explore users’ interests to improve Click-Through Rate (CTR). However, it is very challenging to accurately capture users’ long-term interests from their long-term behavior sequences and give quick responses from the online serving systems. To meet such requirements, existing methods “inadvertently” destroy two basic requirements in long-term sequence modeling:R1) make full use of the entire sequence to keep the information as much as possible;R2) extract information from the most relevant behaviors to keep high relevance between learned interests and current target items. The performance of online serving systems is significantly affected by incomplete and inaccurate user interest information obtained by existing methods. To this end, we propose an efficient two-stage long-term sequence modeling approach, named asEfficieNtClustering based twO-stage interest moDEling (ENCODE), consisting of offline extraction stage and online inference stage. It not only meets the aforementioned two basic requirements but also achieves a desirable balance between online service efficiency and precision. Specifically, in the offline extraction stage, ENCODE clusters the entire behavior sequence and extracts accurate interests. To reduce the overhead of the clustering process, we design a metric learning-based dimension reduction algorithm that preserves the relative pairwise distances of behaviors in the new feature space. While in the online inference stage, ENCODE takes the off-the-shelf user interests to predict the associations with target items. Besides, to further ensure the relevance between user interests and target items, we adopt the same relevance metric throughout the whole pipeline of ENCODE. The extensive experiment and comparison with SOTA on both industrial and public datasets have demonstrated the effectiveness and efficiency of our proposed ENCODE.
Yuhang Zheng 0003, Yinfu Feng, Yunan Ye, Rong Xiao 0005, Long Chen 0016, Xiaosong Yang, Jun Xiao 0001
IEEE Trans. Knowl. Data Eng.8
2025 Decomposed Prototype Learning for Few-Shot Scene Graph Generation
abstract
Today's scene graph generation (SGG) models typically require abundant manual annotations to learn new predicate types. Therefore, it is difficult to apply them to real-world applications with massive uncommon predicate categories whose annotations are hard to collect. In this article, we focus on Few-Shot SGG (FSSGG) , which encourages SGG models to be able to quickly transfer previous knowledge and recognize unseen predicates well with only a few examples. However, current methods for FSSGG are hindered by the high intra-class variance of predicate categories in SGG: On one hand, each predicate category commonly has multiple semantic meanings under different contexts. On the other hand, the visual appearance of relation triplets with the same predicate differs greatly under different subject–object compositions. Such great variance of inputs makes it hard to learn generalizable representation for each predicate category with current few-shot learning (FSL) methods. However, we found that this intra-class variance of predicates is highly related to the composed subjects and objects. To model the intra-class variance of predicates with subject–object context, we propose a novel Decomposed Prototype Learning (DPL) model for FSSGG. Specifically, we first construct a decomposable prototype space to capture diverse semantics and visual patterns of subjects and objects for predicates by decomposing them into multiple prototypes. Afterwards, we integrate these prototypes with different weights to generate query-adaptive predicate representation with more reliable semantics for each query sample. We conduct extensive experiments and compare with various baseline methods to show the effectiveness of our method.
Jun Xiao 0001, Guikun Chen, Yinfu Feng, Yi Yang 0001, Anan Liu, Long Chen 0016
ACM Trans. Multim. Comput. Commun. Appl.2
2024 CoreRec: A Counterfactual Correlation Inference for Next Set Recommendation
abstract
Next set recommendation aims to predict the items that are likely to be bought in the next purchase. Central to this endeavor is the task of capturing intra-set and cross-set correlations among items. However, the modeling of cross-set correlations poses challenges due to specific issues. Primarily, these correlations are often implicit, and the prevailing approach of establishing an indiscriminate link across the entire set of objects neglects factors like purchase frequency and correlations between purchased items. Such hastily formed connections across sets introduce substantial noise. Additionally, the preeminence of high-frequency items in numerous sets could potentially overshadow and distort correlation modeling with respect to low-frequency items. Thus, we devoted to mitigating misleading inter-set correlations. With a fresh perspective rooted in causality, we delve into the question of whether correlations between a particular item and items from other sets should be relied upon for item representation learning and set prediction. Technically, we introduce the Counterfactual Correlation Inference framework for next set recommendation, denoted as CoreRec. This framework establishes a counterfactual scenario in which the recommendation model impedes cross-set correlations to generate intervened predictions. By contrasting these intervened predictions with the original ones, we gauge the causal impact of inter-set neighbors on set prediction—essentially assessing whether they contribute to spurious correlations. During testing, we introduce a post-trained switch module that selects between set-aware item representations derived from either the original or the counterfactual scenarios. To validate our approach, we extensively experiment using three real-world datasets, affirming both the effectiveness of CoreRec and the cogency of our analytical approach.
Chengjiang Long, Shengyu Zhang 0001, Xudong Tang, Zhichao Zhai, Kun Kuang 0001, Jun Xiao 0001
AAAI7
2024 Existence Is Chaos: Enhancing 3D Human Motion Prediction with Uncertainty Consideration
abstract
Human motion prediction is consisting in forecasting future body poses from historically observed sequences. It is a longstanding challenge due to motion's complex dynamics and uncertainty. Existing methods focus on building up complicated neural networks to model the motion dynamics. The predicted results are required to be strictly similar to the training samples with L2 loss in current training pipeline. However, little attention has been paid to the uncertainty property which is crucial to the prediction task. We argue that the recorded motion in training data could be an observation of possible future, rather than a predetermined result. In addition, existing works calculate the predicted error on each future frame equally during training, while recent work indicated that different frames could play different roles. In this work, a novel computationally efficient encoder-decoder model with uncertainty consideration is proposed, which could learn proper characteristics for future frames by a dynamic function. Experimental results on benchmark datasets demonstrate that our uncertainty consideration approach has obvious advantages both in quantity and quality. Moreover, the proposed method could produce motion sequences with much better quality that avoids the intractable shaking artefacts. We believe our work could provide a novel perspective to consider the uncertainty quality for the general motion prediction task and encourage the studies in this field. The code will be available in https://github.com/Motionpre/Adaptive-Salient-Loss-SAGGB.
Ningyu Zhang 0001, Xiaosong Yang, Jun Xiao 0001
AAAI5
2024 Distributionally Generative Augmentation for Fair Facial Attribute Classification
abstract
Facial Attribute Classification (FAC) holds substantial promise in widespread applications. However, FAC models trained by traditional methodologies can be unfair by exhibiting accuracy inconsistencies across varied data sub-populations. This unfairness is largely attributed to bias in data, where some spurious attributes (e.g., Male) statistically correlate with the target attribute (e.g., Smiling). Most of existing fairness-aware methods rely on the labels of spurious attributes, which may be unavailable in practice. This work proposes a novel, generation-based two-stage framework to train a fair FAC model on biased data without additional annotation. Initially, we identify the potential spurious attributes based on generative models. Notably, it enhances interpretability by explicitly showing the spurious attributes in image space. Following this, for each image, we first edit the spurious attributes with a random degree sampled from a uniform distribution, while keeping target attribute unchanged. Then we train a fair FAC model by fostering model invariance to these augmentation. Extensive experiments on three common datasets demonstrate the effectiveness of our method in promoting fairness in FAC without compromising accuracy. Codes are in https://github.com/heqianpei/DiGA.
Fengda Zhang, Qianpei He, Kun Kuang 0001, Long Chen 0016, Chao Wu 0001, Jun Xiao 0001, Hanwang Zhang
CVPR7
2024 DECap: Towards Generalized Explicit Caption Editing via Diffusion Mechanism
Zhen Wang 0004, Xinyun Jiang, Jun Xiao 0001, Long Chen 0016
ECCV (43)3
2024 Seeing Beyond Classes: Zero-Shot Grounded Situation Recognition via Language Explainer
abstract
Benefiting from strong generalization ability, pre-trained vision-language models (VLMs), e.g., CLIP, have been widely utilized in zero-shot scene understanding. Unlike simple recognition tasks, grounded situation recognition (GSR) requires the model not only to classify salient activity (verb) in the image, but also to detect all semantic roles that participate in the action. This complex task usually involves three steps: verb recognition, semantic role grounding, and noun recognition. Directly employing class-based prompts with VLMs and grounding models for this task suffers from several limitations, e.g., it struggles to distinguish ambiguous verb concepts, accurately localize roles with fixed verb-centric template input, and achieve context-aware noun predictions. In this paper, we argue that these limitations stem from the model's poor understanding of verb/noun classes. To this end, we introduce a new approach for zero-shot GSR via Language EXplainer (LEX), which significantly boosts the model's comprehensive capabilities through three explainers: 1) verb explainer, which generates general verb-centric descriptions to enhance the discriminability of different verb classes; 2) grounding explainer, which rephrases verb-centric templates for clearer understanding, thereby enhancing precise semantic role localization; and 3) noun explainer, which creates scene-specific noun descriptions to ensure context-aware noun recognition. By equipping each step of the GSR process with an auxiliary explainer, LEX facilitates complex scene understanding in real-world scenarios. Our extensive validations on the SWiG dataset demonstrate LEX's effectiveness and interoperability in zero-shot GSR.
Jiaming Lei, Lin Li 0065, Chunping Wang 0001, Jun Xiao 0001, Long Chen 0016
ACM Multimedia4
2024 Neural Interaction Energy for Multi-Agent Trajectory Prediction
abstract
Maintaining temporal stability is crucial in multi-agent trajectory prediction. Insufficient regularization to uphold this temporal stability often results in fluctuations in kinematic states, leading to inconsistent predictions and the amplification of errors. In this study, we introduce a framework called Multi-Agent Trajectory prediction via neural interaction Energy (MATE). This framework assesses the interactive motion of agents by employing neural interaction energy, which captures the dynamics of interactions and illustrates their influence on the future trajectories of agents. To bolster temporal stability, we introduce two constraints: inter-agent interaction constraint and intra-agent motion constraint. These constraints work together to ensure temporal stability at both the system and agent levels, effectively mitigating prediction fluctuations inherent in multi-agent systems. Comparative evaluations against previous methods on four diverse datasets, including simulated and real-world scenarios, highlight the superior prediction accuracy and generalization capabilities of our model.
Kaixin Shen, Ruijie Quan, Linchao Zhu, Jun Xiao 0001, Yi Yang 0001
ACM Multimedia4
2024 $\text{Di}^2\text{Pose}$: Discrete Diffusion Model for Occluded 3D Human Pose Estimation
abstract
Diffusion models have demonstrated their effectiveness in addressing the inherent uncertainty and indeterminacy in monocular 3D human pose estimation (HPE). Despite their strengths, the need for large search spaces and the corresponding demand for substantial training data make these models prone to generating biomechanically unrealistic poses. This challenge is particularly noticeable in occlusion scenarios, where the complexity of inferring 3D structures from 2D images intensifies. In response to these limitations, we introduce the **Di**screte **Di**ffusion **Pose** (**$\text{Di}^2\text{Pose}$**), a novel framework designed for occluded 3D HPE that capitalizes on the benefits of a discrete diffusion model. Specifically, **$\text{Di}^2\text{Pose}$** employs a two-stage process: it first converts 3D poses into a discrete representation through a pose quantization step, which is subsequently modeled in latent space through a discrete diffusion process. This methodological innovation restrictively confines the search space towards physically viable configurations and enhances the model’s capability to comprehend how occlusions affect human pose within the latent space. Extensive evaluations conducted on various benchmarks (e.g., Human3.6M, 3DPW, and 3DPW-Occ) have demonstrated its effectiveness.
Jun Xiao 0001, Chunping Wang 0001, Wei Liu 0005, Long Chen 0016
NeurIPS2
2024 UN-η: An offline adaptive normalization method for deploying transformers
Zheyang Li, Kai Zhang 0055, Chaoxiang Lan, Huanlong Zhang, Wenming Tan, Jun Xiao 0001, Shiliang Pu
Knowl. Based Syst.7
2024 NICEST: Noisy Label Correction and Training for Robust Scene Graph Generation
abstract
Nearly all existing scene graph generation (SGG) models have overlooked the ground-truth annotation qualities of mainstream SGG datasets, i.e., they assume: 1) all the manually annotated positive samples are equally correct; 2) all the un-annotated negative samples are absolutely background. In this paper, we argue that neither of the assumptions applies to SGG: there are numerous “noisy” ground-truth predicate labels that break these two assumptions and harm the training of unbiased SGG models. To this end, we propose a novelNoIsy labelCorrEction andSampleTrainingstrategy for SGG:NICEST, which rules out these noisy label issues by generating high-quality samples and designing an effective training strategy. Specifically, it consists of: 1)NICE: it detects noisy samples and then reassigns higher-quality soft predicate labels to them. To achieve this goal, NICE contains three main steps: negative Noisy Sample Detection (Neg-NSD), positive NSD (Pos-NSD), and Noisy Sample Correction (NSC). Firstly, in Neg-NSD, it is treated as an out-of-distribution detection problem, and the pseudo labels are assigned to all detected noisy negative samples. Then, in Pos-NSD, we use a density-based clustering algorithm to detect noisy positive samples. Lastly, in NSC, we use weighted KNN to reassign more robust soft predicate labels rather than hard labels to all noisy positive samples. 2)NIST: it is a multi-teacher knowledge distillation based training strategy, which enables the model to learn unbiased fusion knowledge. A dynamic trade-off weighting strategy in NIST is designed to penalize the bias of different teachers. Due to the model-agnostic nature of both NICE and NIST, NICEST can be seamlessly incorporated into any SGG architecture to boost its performance on different predicate categories. In addition, to better assess the generalization ability of SGG models, we propose a new benchmark,VG-OOD, by reorganizing the prevalent VG dataset. This reorganization deliberately makes the predicate distributions between the training and test sets as different as possible for each subject-object category pair. This new benchmark helps disentangle the influence of subject-object category biases. Extensive ablations and results on different backbones and tasks have attested to the effectiveness and generalization ability of each component of NICEST.
Lin Li 0065, Jun Xiao 0001, Hanrong Shi, Hanwang Zhang, Yi Yang 0001, Wei Liu 0005, Long Chen 0016
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 IDPro: Flexible Interactive Video Object Segmentation by ID-Queried Concurrent Propagation
abstract
Interactive Video Object Segmentation (iVOS) is inherently demanding, requiring real-time interaction between humans and computers. Enhancing user experience involves considerations such as user input habits, segmentation quality, running time, and memory consumption. However, existing methods compromise user experience by employing a single input mode and exhibiting slow running speeds. Specifically, these approaches restrict user interaction to a single frame, limiting the expression of user intent. To overcome these limitations and better align with user habits, we introduce a framework that facilitates flexible input modes by ID-queried concurrent propagation (IDPro). In particular, we have devised the Across-Frame Interaction Module (AFI), allowing users to freely annotate various objects across multiple frames. The AFI module transfers scribble information across interactive frames, generating multi-frame masks. Additionally, we leverage an id-queried mechanism to process multiple objects. To achieve more efficient propagation and a lightweight model, we propose a truncated re-propagation strategy, replacing the previous multi-round fusion module, which employs an across-round memory that stores crucial interaction information. Our SwinB-IDPro attains a new state-of-the-art performance on DAVIS 2017 (89.6%,${\mathcal {J}}\& {\mathcal {F}}\text{@}60$). Furthermore, our R50-IDPro exhibits over${3 \times }$faster performance than the leading competitor in challenging multi-object scenarios.
Tao Jiang 0042, Zongxin Yang, Yi Yang 0001, Yueting Zhuang, Jun Xiao 0001
IEEE Trans. Circuits Syst. Video Technol.6
2024 Label Semantic Knowledge Distillation for Unbiased Scene Graph Generation
abstract
The Scene Graph Generation (SGG) task aims to detect all the objects and their pairwise visual relationships in a given image. Although SGG has achieved remarkable progress over the last few years, almost all existing SGG models follow the same training paradigm: they treat both object and predicate classification in SGG as a single-label classification problem, and the ground-truths are one-hot target labels. However, this prevalent training paradigm has overlooked two characteristics of current SGG datasets: 1) For positive samples, some specific subject-object instances may have multiple reasonable predicates. 2) For negative samples, there are numerous missing annotations. Regardless of the two characteristics, SGG models are easy to be confused and make wrong predictions. To this end, we propose a novel model-agnostic Label Semantic Knowledge Distillation (LS-KD) for unbiased SGG. Specifically, LS-KD dynamically generates a “soft” label for each subject-object instance by fusing a predicted Label Semantic Distribution (LSD) with its original one-hot target label. LSD reflects the correlations between this instance and multiple predicate categories. Meanwhile, we propose two different strategies to predict LSD: iterative self-KD and synchronous self-KD. Extensive ablations and results on three SGG tasks have attested to the superiority and generality of our proposed LS-KD, which can consistently achieve decent trade-off performance between different predicate categories.
Lin Li 0065, Jun Xiao 0001, Hanrong Shi, Wenxiao Wang 0001, Jian Shao 0001, Anan Liu, Yi Yang 0001, Long Chen 0016
IEEE Trans. Circuits Syst. Video Technol.2
2024 Knowledge-Guided Causal Intervention for Weakly-Supervised Object Localization
abstract
Previous weakly-supervised object localization (WSOL) methods aim to expand activation map discriminative areas to cover the whole objects, yet neglect two inherent challenges when relying solely on image-level labels. First, the “entangled context” issue arises from object-context co-occurrence (e.g., fish and water), making the model inspection hard to distinguish object boundaries clearly. Second, the “C-L dilemma” issue results from the information decay caused by the pooling layers, which struggle to retain both the semantic information for precise classification and those essential details for accurate localization, leading to a trade-off in performance. In this paper, we propose a knowledge-guided causal intervention method, dubbed KG-CI-CAM, to address these two under-explored issues in one go. More specifically, we tackle the co-occurrence context confounder problem via causal intervention, which explores the causalities among image features, contexts, and categories to eliminate the biased object-context entanglement in the class activation maps. Based on the disentangled object feature, we introduce a multi-source knowledge guidance framework to strike a balance between absorbing classification knowledge and localization knowledge during model training. Extensive experiments conducted on several benchmark datasets demonstrate the effectiveness of KG-CI-CAM in learning distinct object boundaries amidst confounding contexts and mitigating the dilemma between classification and localization performance.
Feifei Shao, Yawei Luo, Fei Gao 0014, Yi Yang 0001, Jun Xiao 0001
IEEE Trans. Knowl. Data Eng.5
2024 Taking a Closer Look At Visual Relation: Unbiased Video Scene Graph Generation With Decoupled Label Learning
abstract
Current video-based scene graph generation (VidSGG) methods have been found to perform poorly in predicting predicates that are less represented due to the inherently biased distribution of the training data. In this paper, we take a closer look at the inherent characteristics of predicates and identify that most visual relations (e.g.sit_above) involve both actional pattern (sit) and spatial pattern (above), while the distribution bias is much less severe at the pattern level. Based on this insight, we propose a decoupled label learning (DLL) paradigm to address the intractable visual relation prediction from the pattern-level perspective. Specifically, DLL decouples the predicate labels and adopts separate classifiers to learn actional and spatial patterns respectively. The patterns are then combined and mapped back to the predicate. Moreover, we propose a knowledge-level label decoupling method to transfer non-target knowledge from head predicates to tail predicates within the same pattern to calibrate the distribution of tail classes. We validate the effectiveness of DLL on the commonly used VidSGG benchmark, i.e. VidVRD. Extensive experiments demonstrate that the DLL offers a remarkably simple but highly effective solution to the long-tailed problem, achieving the state-of-the-art VidSGG performance.
Yawei Luo, Zhiqing Chen, Tao Jiang 0042, Yi Yang 0001, Jun Xiao 0001
IEEE Trans. Multim.6
2024 Improving Reference-Based Distinctive Image Captioning with Contrastive Rewards
abstract
Distinctive Image Captioning (DIC)—generating distinctive captions that describe the unique details of a target image—has received considerable attention over the last few years. A recent DIC method proposes to generate distinctive captions by comparing the target image with a set of semantic-similar reference images, i.e., reference-Based DIC (Ref-DIC). It aims to force the generated captions to distinguish between the target image and the reference image. Unfortunately, reference images used by existing Ref-DIC works are easy to distinguish: these reference images only resemble the target image at scene-level and have few common objects, such that a Ref-DIC model can trivially generate distinctive captions even without considering the reference images. For example, if the target image contains objects “ towel ” and “ toilet ” while all reference images are without them, then a simple caption “ A bathroom with a towel and a toilet ” is distinctive enough to tell apart target and reference images. To ensure Ref-DIC models really perceive the unique objects (or attributes) in target images, we first propose two new Ref-DIC benchmarks. Specifically, we design a two-stage matching mechanism, which strictly controls the similarity between the target and reference images at the object-/attribute-level (vs. scene-level). Second, to generate distinctive captions, we develop a Transformer-based Ref-DIC baseline TransDIC . It not only extracts visual features from the target image but also encodes the differences between objects in the target and reference images. Taking one step further, we propose a stronger TransDIC \({++}\) , which consists of an extra contrastive learning module to make full use of the reference images. This new module is model-agnostic, which can be easily incorporated into various Ref-DIC architectures. Finally, for more trustworthy benchmarking, we propose a new evaluation metric named DisCIDEr for Ref-DIC, which evaluates both the accuracy and distinctiveness of the generated captions. Experimental results demonstrate that our TransDIC \({++}\) can generate distinctive captions. Besides, it outperforms several state-of-the-art models on the two new benchmarks over different metrics.
Yangjun Mao, Jun Xiao 0001, Meng Cao 0002, Jian Shao 0001, Yueting Zhuang, Long Chen 0016
ACM Trans. Multim. Comput. Commun. Appl.2
2023 Bit-shrinking: Limiting Instantaneous Sharpness for Improving Post-training Quantization
abstract
Post-training quantization (PTQ) is an effective compression method to reduce the model size and computational cost. However, quantizing a model into a low-bit one, e.g., lower than 4, is difficult and often results in non-negligible performance degradation. To address this, we investigate the loss landscapes of quantized networks with various bit-widths. We show that the network with more ragged loss surface, is more easily trapped into bad local minima, which mostly appears in low-bit quantization. A deeper analysis indicates, the ragged surface is caused by the injection of excessive quantization noise. To this end, we detach a sharpness term from the loss which reflects the impact of quantization noise. To smooth the rugged loss surface, we propose to limit the sharpness term small and stable during optimization. Instead of directly optimizing the target bit network, we design a self-adapted shrinking scheduler for the bit-width in continuous domain from high bit-width to the target by limiting the increasing sharpness term within a proper range. It can be viewed as iteratively adding small “instant” quantization noise and adjusting the network to eliminate its impact. Widely experiments including classification and detection tasks demonstrate the effectiveness of the Bit-shrinking strategy in PTQ. On the Vision Transformer models, our INT8 and INT6 models drop within 0.5% and 1.5% Top-1 accuracy, respectively. On the traditional CNN networks, our INT4 quantized models drop within 1.3% and 3.5% Top-1 accuracy on ResNet18 and MobileNetV2 without fine-tuning, which achieves the state-of-the-art performance.
Zheyang Li, Wenming Tan, Ye Ren, Jun Xiao 0001, Shiliang Pu
CVPR6
2023 Compositional Feature Augmentation for Unbiased Scene Graph Generation
abstract
Scene Graph Generation (SGG) aims to detect all the visual relation tripletsin a given image. With the emergence of various advanced techniques for better utilizing both the intrinsic and extrinsic information in each relation triplet, SGG has achieved great progress over the recent years. However, due to the ubiquitous long-tailed predicate distributions, today’s SGG models are still easily biased to the head predicates. Currently, the most prevalent debiasing solutions for SGG are re-balancing methods, e.g., changing the distributions of original training samples. In this paper, we argue that all existing re-balancing strategies fail to increase the diversity of the relation triplet features of each predicate, which is critical for robust SGG. To this end, we propose a novel Compositional Feature Augmentation (CFA) strategy, which is the first unbiased SGG work to mitigate the bias issue from the perspective of increasing the diversity of triplet features. Specifically, we first decompose each relation triplet feature into two components: intrinsic feature and extrinsic feature, which correspond to the intrinsic characteristics and extrinsic contexts of a relation triplet, respectively. Then, we design two different feature augmentation modules to enrich the feature diversity of original relation triplets by replacing or mixing up either their intrinsic or extrinsic features from other samples. Due to its model-agnostic nature, CFA can be seamlessly incorporated into various SGG frameworks. Extensive ablations have shown that CFA achieves a new state-of-the-art performance on the trade-off between different metrics.
Lin Li 0065, Guikun Chen, Jun Xiao 0001, Yi Yang 0001, Chunping Wang 0001, Long Chen 0016
ICCV3
2023 Video Scene Graph Generation from Single-Frame Weak Supervision
Jun Xiao 0001, Long Chen 0016
ICLR2
2023 Compositional Prompt Tuning with Motion Cues for Open-vocabulary Video Relation Detection
Kaifeng Gao, Long Chen 0016, Hanwang Zhang, Jun Xiao 0001, Qianru Sun
ICLR4
2023 Fairness-aware Contrastive Learning with Partially Annotated Sensitive Attributes
Fengda Zhang, Kun Kuang 0001, Long Chen 0016, Chao Wu 0001, Jun Xiao 0001
ICLR6
2023 Addressing Predicate Overlap in Scene Graph Generation with Semantic Granularity Controller
abstract
Semantic overlap between predicates (e.g., riding versus on) occurs inevitably when describing a scene. However, most existing Scene Graph Generation (SGG) works sidestep it by modeling the semantic overlap at category-level and assigning merely one-hot target to each sample, which hurt the performance on other reasonable predicates. In this paper, we argue that semantic overlap between predicates tends to vary in different abstract patterns, and a subject-object pair should retain multiple reasonable predicates. To this end, we make an early attempt to reformulate SGG as a partial multi-label learning problem and accordingly propose a model-agnostic Semantic Granularity Controller (SGC). SGC consists of a pattern-specific controller, partial multi-label learning, and controllable inference. The former two solve semantic confusion during training, while the latter makes the semantic granularity of prediction controllable. Extensive experiments demonstrate that SGC can improve the performance of SGG and guide the model to predict coarse/fine-grained predicates.
Guikun Chen, Lin Li 0065, Yawei Luo, Jun Xiao 0001
ICME4
2023 Dark Knowledge Balance Learning for Unbiased Scene Graph Generation
abstract
One of the major obstacles that hinders the current scene graph generation (SGG) performance lies in the severe predicate annotation bias. Conventional solutions to this problem are mainly based on reweighting/resampling heuristics. Despite achieving some improvements on tail classes, these methods are prone to cause serious performance degradation of head predicates. In this paper, we propose to tackle this problem from a brand-new perspective of dark knowledge. In consideration of the unique nature of SGG that requires a large number of negative samples to be employed for predicate learning, we design to capitalize on the dark knowledge contained in negative samples for debiasing the predicate distribution. Along such vein, we propose a novel SGG method dubbed Dark Knowledge Balance Learning (DKBL). In DKBL, we first design a dark knowledge balancing loss, which helps the model learn to balance head and tail predicates while maintaining the overall performance. We further introduce a dark knowledge semantic enhancement module to better encode the semantics of predicates. DKBL is orthogonal to existing SGG methods and can be easily plugged into their training process for further improvement. Extensive experiments on VG dataset show that the proposed DKBL can consistently achieve well trade-off performance between head and tail predicates, which is significantly better than previous state-of-the-art methods. The code is available in https://github.com/chenzqing/DKBL.
Zhiqing Chen, Yawei Luo, Jian Shao 0001, Yi Yang 0001, Chunping Wang 0001, Lei Chen 0082, Jun Xiao 0001
ACM Multimedia7
2023 FedAA: Using Non-sensitive Modalities to Improve Federated Learning while Preserving Image Privacy
abstract
Federated learning aims to train a better global model without sharing the sensitive training samples (usually images) of local clients. Since the sample distributions in local clients tend to be different from each other (i.e., non-IID), one of the major challenges for federated learning is to alleviate model degradation when aggregating local models. The degradation can be attributed to the weight divergence that quantifies the difference of local models from different training processes. Furthermore, non-IID also results in feature space heterogeneity during local training, making neurons of local models in the same location have different functions and further exacerbating weight divergence. In this paper, we demonstrate that the problem can be solved by sharing information from the non-sensitive modality (e.g., metadata, non-sensitive descriptions, etc.) while keeping the sensitive information of images protected. In particular, we propose Federated Learning with Adversarial Example and Adversarial Identifier (FedAA) that trains adversarial examples based on the shared non-sensitive modality to fine-tune local models before global aggregation. The training of local models is enhanced by client identifiers that discriminate the source of inputs to force different local models to get similar outputs and be more homogeneous during the local training. Experiments show that FedAA significantly outperforms recent non-IID federated learning algorithms while preserving image privac, by sharing information from non-sensitive modalities.
Dong Chen 0017, Siliang Tang, Zijin Shen, Guoming Wang, Jun Xiao 0001, Yueting Zhuang, Carl Yang 0001
ACM Multimedia5
2023 CATR: Combinatorial-Dependence Audio-Queried Transformer for Audio-Visual Video Segmentation
abstract
Audio-visual video segmentation (AVVS) aims to generate pixel-level maps of sound-producing objects within image frames and ensure the maps faithfully adheres to the given audio, such as identifying and segmenting a singing person in a video. However, existing methods exhibit two limitations: 1) they address video temporal features and audio-visual interactive features separately, disregarding the inherent spatial-temporal dependence of combined audio and video, and 2) they inadequately introduce audio constraints and object-level information during the decoding stage, resulting in segmentation outcomes that fail to comply with audio directives. To tackle these issues, we propose a decoupled audio-video transformer that combines audio and video features from their respective temporal and spatial dimensions, capturing their combined dependence. To optimize memory consumption, we design a block, which, when stacked, enables capturing audio-visual fine-grained combinatorial-dependence in a memory-efficient manner. Additionally, we introduce audio-constrained queries during the decoding phase. These queries contain rich object-level information, ensuring the decoded mask adheres to the sounds. Experimental results confirm our approach's effectiveness, with our framework achieving a new SOTA performance on all three datasets using two backbones. The code is available at https://github.com/aspirinone/CATR.github.io.
Zongxin Yang, Lei Chen 0082, Yi Yang 0001, Jun Xiao 0001
ACM Multimedia5
2023 Triple Correlations-Guided Label Supplementation for Unbiased Video Scene Graph Generation
abstract
Video-based scene graph generation (VidSGG) is an approach that aims to represent video content in a dynamic graph by identifying visual entities and their relationships. Due to the inherently biased distribution and missing annotations in the training data, current VidSGG methods have been found to perform poorly on less-represented predicates. In this paper, we propose an explicit solution to address this under-explored issue by supplementing missing predicates that should be included in the ground-truth annotations. Dubbed Trico, our method seeks to supplement the missing predicates that are supposed to appear in the ground-truth annotations, by exploring three complementary spatio-temporal correlations. Guided by these correlations, the missing labels can be effectively supplemented thus achieving an unbiased predicate predictions. We validate the effectiveness of Trico on the most widely used VidSGG datasets, i.e., VidVRD and VidOR. Extensive experiments demonstrate the state-of-the-art performance achieved by Trico, particularly on those tail predicates. The code is available in the supplementary material.
Kaifeng Gao, Yawei Luo, Tao Jiang 0042, Fei Gao 0014, Jian Shao 0001, Jun Xiao 0001
ACM Multimedia8
2023 Generalized Universal Domain Adaptation with Generative Flow Networks
abstract
We introduce a new problem in unsupervised domain adaptation, termed as Generalized Universal Domain Adaptation (GUDA), which aims to achieve precise prediction of all target labels including unknown categories. GUDA bridges the gap between label distribution shift-based and label space mismatch-based variants, essentially categorizing them as a unified problem, guiding to a comprehensive framework for thoroughly solving all the variants. The key challenge of GUDA is developing and identifying novel target categories while estimating the target label distribution. To address this problem, we take advantage of the powerful exploration capability of generative flow networks and propose an active domain adaptation algorithm named GFlowDA, which selects diverse samples with probabilities proportional to a reward function. To enhance the exploration capability and effectively perceive the target label distribution, we tailor the states and rewards, and introduce an efficient solution for parent exploration and state transition. We also propose a training paradigm for GUDA called Generalized Universal Adversarial Network (GUAN), which involves collaborative optimization between GUAN and GFlowNet. Theoretical analysis highlights the importance of exploration, and extensive experiments on benchmark datasets demonstrate the superiority of GFlowDA.
Didi Zhu, Yinchuan Li, Yunfeng Shao 0001, Jianye Hao, Fei Wu 0001, Kun Kuang 0001, Jun Xiao 0001, Chao Wu 0001
ACM Multimedia7
2023 Two Heads are Better Than One: A Simple Exploration Framework for Efficient Multi-Agent Reinforcement Learning
abstract
Exploration strategy plays an important role in reinforcement learning, especially in sparse-reward tasks. In cooperative multi-agent reinforcement learning~(MARL), designing a suitable exploration strategy is much more challenging due to the large state space and the complex interaction among agents. Currently, mainstream exploration methods in MARL either contribute to exploring the unfamiliar states which are large and sparse, or measuring the interaction among agents with high computational costs. We found an interesting phenomenon that different kinds of exploration plays a different role in different MARL scenarios, and choosing a suitable one is often more effective than designing an exquisite algorithm. In this paper, we propose a exploration method that incorporate the \underline{C}uri\underline{O}sity-based and \underline{IN}fluence-based exploration~(COIN) which is simple but effective in various situations. First, COIN measures the influence of each agent on the other agents based on mutual information theory and designs it as intrinsic rewards which are applied to each individual value function. Moreover, COIN computes the curiosity-based intrinsic rewards via prediction errors which are added to the extrinsic reward. For integrating the two kinds of intrinsic rewards, COIN utilizes a novel framework in which they complement each other and lead to a sufficient and effective exploration on cooperative MARL tasks. We perform extensive experiments on different challenging benchmarks, and results across different scenarios show the superiority of our method.
Jiahui Li 0003, Kun Kuang 0001, Baoxiang Wang 0001, Fei Wu 0001, Jun Xiao 0001, Long Chen 0016
NeurIPS6
2023 Zero-shot Visual Relation Detection via Composite Visual Cues from Large Language Models
abstract
Pretrained vision-language models, such as CLIP, have demonstrated strong generalization capabilities, making them promising tools in the realm of zero-shot visual recognition. Visual relation detection (VRD) is a typical task that identifies relationship (or interaction) types between object pairs within an image. However, naively utilizing CLIP with prevalent class-based prompts for zero-shot VRD has several weaknesses, e.g., it struggles to distinguish between different fine-grained relation types and it neglects essential spatial information of two objects. To this end, we propose a novel method for zero-shot VRD: RECODE, which solves RElation detection via COmposite DEscription prompts. Specifically, RECODE first decomposes each predicate category into subject, object, and spatial components. Then, it leverages large language models (LLMs) to generate description-based prompts (or visual cues) for each component. Different visual cues enhance the discriminability of similar relation categories from different perspectives, which significantly boosts performance in VRD. To dynamically fuse different cues, we further introduce a chain-of-thought method that prompts LLMs to generate reasonable weights for different visual cues. Extensive experiments on four VRD benchmarks have demonstrated the effectiveness and interpretability of RECODE.
Lin Li 0065, Jun Xiao 0001, Guikun Chen, Jian Shao 0001, Yueting Zhuang, Long Chen 0016
NeurIPS2
2023 Differentiated matching for individual and average treatment effect estimation
Ziyu Zhao 0001, Kun Kuang 0001, Bo Li 0064, Peng Cui 0001, Runze Wu 0001, Jun Xiao 0001, Fei Wu 0001
Data Min. Knowl. Discov.6
2023 Question-guided feature pyramid network for medical visual question answering
Yonglin Yu, Hanrong Shi, Lin Li 0065, Jun Xiao 0001
Expert Syst. Appl.5
2023 Unsupervised self-training correction learning for 2D image-based 3D model retrieval
Yaqian Zhou 0002, Yu Liu 0004, Jun Xiao 0001, Min Liu 0008, Xuanya Li, Anan Liu
Inf. Process. Manag.3
2023 Federated unsupervised representation learning
abstract
To leverage the enormous amount of unlabeled data on distributed edge devices, we formulate a new problem in federated learning called federated unsupervised representation learning (FURL) to learn a common representation model without supervision while preserving data privacy. FURL poses two new challenges: (1) data distribution shift (non-independent and identically distributed, non-IID) among clients would make local models focus on different categories, leading to the inconsistency of representation spaces; (2) without unified information among the clients in FURL, the representations across clients would be misaligned. To address these challenges, we propose the federated contrastive averaging with dictionary and alignment (FedCA) algorithm. FedCA is composed of two key modules: a dictionary module to aggregate the representations of samples from each client which can be shared with all clients for consistency of representation space and an alignment module to align the representation of each client on a base model trained on public data. We adopt the contrastive approach for local model training. Through extensive experiments with three evaluation protocols in IID and non-IID settings, we demonstrate that FedCA outperforms all baselines with significant margins.
Fengda Zhang, Kun Kuang 0001, Long Chen 0016, Zhaoyang You, Tao Shen 0002, Jun Xiao 0001, Yin Zhang 0006, Chao Wu 0001, Fei Wu 0001, Yueting Zhuang
Frontiers Inf. Technol. Electron. Eng.6
2023 Counterfactual Samples Synthesizing and Training for Robust Visual Question Answering
abstract
Today's VQA models still tend to capture superficial linguistic correlations in the training set and fail to generalize to the test set with different QA distributions. To reduce these language biases, recent VQA works introduce an auxiliary question-only model to regularize the training of targeted VQA model, and achieve dominating performance on diagnostic benchmarks for out-of-distribution testing. However, due to the complex model design, ensemble-based methods are unable to equip themselves with two indispensable characteristics of an ideal VQA model: 1) Visual-explainable: The model should rely on the right visual regions when making decisions. 2) Question-sensitive: The model should be sensitive to the linguistic variations in questions. To this end, we propose a novel model-agnostic Counterfactual Samples Synthesizing and Training (CSST) strategy. After training with CSST, VQA models are forced to focus on all critical objects and words, which significantly improves both visual-explainable and question-sensitive abilities. Specifically, CSST is composed of two parts: Counterfactual Samples Synthesizing (CSS) and Counterfactual Samples Training (CST). CSS generates counterfactual samples by carefully masking critical objects in images or words in questions and assigning pseudo ground-truth answers. CST not only trains the VQA models with both complementary samples to predict respective ground-truth answers, but also urges the VQA models to further distinguish the original samples and superficially similar counterfactual ones. To facilitate the CST training, we propose two variants of supervised contrastive loss for VQA, and design an effective positive and negative sample selection mechanism based on CSS. Extensive experiments have shown the effectiveness of CSST. Particularly, by building on top of model LMH+SAR (Clark et al. 2019), (Si et al. 2021), we achieve record-breaking performance on all out-of-distribution benchmarks (e.g., VQA-CP v2, VQA-CP v1, and GQA-OOD).
Long Chen 0016, Yuhang Zheng 0003, Yulei Niu, Hanwang Zhang, Jun Xiao 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Dual-Path Rare Content Enhancement Network for Image and Text Matching
abstract
Image and text matching plays a crucial role in bridging the cross-modal gap between vision and language, and has achieved great progress due to the deep learning. However, the existing methods still suffer from the long-tail problem, where only a small proportion contains highly frequent semantics and a long tail proportion is constructed by rare semantics. In this paper, we propose a novel Dual-path Rare Content Enhancement Network (DRCE) to tackle the long-tail issue. Specifically, the Cross-modal Representation Enhancement (CRE) and Cross-modal Association Enhancement (CAE) are proposed to construct dual-path structure to enhance rare content representation and association with the benefit of cross-modal prior knowledge. This structure can effectively exploit the complementary cross-modal relation from different aspects and fuse these information in an adaptively manner by the proposed Adaptive Fusion Strategy (AFS). Moreover, we also propose an alternative re-ranking strategy (ARR) to explore the reciprocal contextual information to refine image-text matching results, which can further suppress the negative effect of long-tail effect. Extensive experiments on two large-scale datasets show the significant improvements and validate the superiority of our method.
Yan Wang 0114, Yuting Su 0001, Wenhui Li 0001, Jun Xiao 0001, Xuanya Li, Anan Liu
IEEE Trans. Circuits Syst. Video Technol.4
2023 VL-NMS: Breaking Proposal Bottlenecks in Two-stage Visual-language Matching
abstract
The prevailing framework for matching multimodal inputs is based on a two-stage process: (1) detecting proposals with an object detector and (2) matching text queries with proposals. Existing two-stage solutions mostly focus on the matching step. In this article, we argue that these methods overlook an obvious mismatch between the roles of proposals in the two stages: they generate proposals solely based on the detection confidence (i.e., query-agnostic), hoping that the proposals contain all instances mentioned in the text query (i.e., query-aware). Due to this mismatch, chances are that proposals relevant to the text query are suppressed during the filtering process, which in turn bounds the matching performance. To this end, we propose VL-NMS, which is the first method to yield query-aware proposals at the first stage. VL-NMS regards all mentioned instances as critical objects and introduces a lightweight module to predict a score for aligning each proposal with a critical object. These scores can guide the NMS operation to filter out proposals irrelevant to the text query, increasing the recall of critical objects, and resulting in a significantly improved matching performance. Since VL-NMS is agnostic to the matching step, it can be easily integrated into any state-of-the-art two-stage matching method. We validate the effectiveness of VL-NMS on three multimodal matching tasks, namely referring expression grounding, phrase grounding, and image-text matching. Extensive ablation studies on several baselines and benchmarks consistently demonstrate the superiority of VL-NMS.
Chenchi Zhang, Jun Xiao 0001, Hanwang Zhang, Jian Shao 0001, Yueting Zhuang, Long Chen 0016
ACM Trans. Multim. Comput. Commun. Appl.3
2022 Rethinking the Evaluation of Unbiased Scene Graph Generation
Long Chen 0016, Jian Shao 0001, Shaoning Xiao, Songyang Zhang 0004, Jun Xiao 0001
BMVC6
2022 DUDA: Online-Offline Dual Domain Adaption for Semantic Segmentation
An-tao Pan, Yawei Luo, Yi Yang 0001, Jun Xiao 0001
BMVC4
2022 Classification-Then-Grounding: Reformulating Video Scene Graphs as Temporal Bipartite Graphs
abstract
Today's VidSGG models are all proposal-based methods, i.e., they first generate numerous paired subject-object snippets as proposals, and then conduct predicate classification for each proposal. In this paper, we argue that this prevalent proposal-based framework has three inherent drawbacks: 1) The ground-truth predicate labels for proposals are partially correct. 2) They break the high-order relations among different predicate instances of a same subject-object pair. 3) VidSGG performance is upper-bounded by the quality of the proposals. To this end, we propose a new classification-then-grounding framework for VidSGG, which can avoid all the three overlooked drawbacks. Meanwhile, under this framework, we reformulate the video scene graphs as temporal bipartite graphs, where the entities and predicates are two types of nodes with time slots, and the edges denote different semantic roles between these nodes. This formulation takes full advantage of our new framework. Accordingly, we further propose a novel BIpartite Graph based SGG model: BIG. It consists of a classification stage and a grounding stage, where the former aims to classify the categories of all the nodes and the edges, and the latter tries to localize the temporal location of each relation instance. Extensive ablations on two VidSGG datasets have attested to the effectiveness of our framework and BIG. Code is available at https://github.com/Dawn-LX/VidSGG-BIG.
Kaifeng Gao, Long Chen 0016, Yulei Niu, Jian Shao 0001, Jun Xiao 0001
CVPR5
2022 The Devil is in the Labels: Noisy Label Correction for Robust Scene Graph Generation
abstract
Unbiased SGG has achieved significant progress over recent years. However, almost all existing SGG models have overlooked the ground-truth annotation qualities of prevailing SGG datasets, i.e., they always assume: 1) all the manually annotated positive samples are equally correct; 2) all the un-annotated negative samples are absolutely background. In this paper, we argue that both assumptions are inapplicable to SGG: there are numerous “noisy” ground-truth predicate labels that break these two assumptions, and these noisy samples actually harm the training of unbiased SGG models. To this end, we propose a novel model-agnostic NoIsy label CorrEction strategy for SGG: NICE. NICE can not only detect noisy samples but also reassign more high-quality predicate labels to them. After the NICE training, we can obtain a cleaner version of SGG dataset for model training. Specifically, NICE consists of three components: negative Noisy Sample Detection (Neg-NSD), positive NSD (Pos-NSD), and Noisy Sample Correction (NSC). Firstly, in Neg-NSD, we formulate this task as an out-of-distribution detection problem, and assign pseudo labels to all detected noisy negative samples. Then, in Pos-NSD, we use a clustering-based algorithm to divide all positive samples into multiple sets, and treat the samples in the noisiest set as noisy positive samples. Lastly, in NSC, we use a simple but effective weighted KNN to reassign new predicate labels to noisy positive samples. Extensive results on different backbones and tasks have attested to the effectiveness and generalization abilities of each component of NICE.
Lin Li 0065, Long Chen 0016, Songyang Zhang 0004, Jun Xiao 0001
CVPR6
2022 Rethinking Data Augmentation for Robust Visual Question Answering
Long Chen 0016, Yuhang Zheng 0003, Jun Xiao 0001
ECCV (36)3
2022 Explicit Image Caption Editing
Zhen Wang 0004, Long Chen 0016, Guangxing Han, Yulei Niu, Jian Shao 0001, Jun Xiao 0001
ECCV (36)7
2022 Rethinking Multi-Modal Alignment in Multi-Choice VideoQA from Feature and Sample Perspectives
abstract
Reasoning about causal and temporal event relations in videos is a new destination of Video Question Answering (VideoQA).The major stumbling block to achieve this purpose is the semantic gap between language and video since they are at different levels of abstraction.Existing efforts mainly focus on designing sophisticated architectures while utilizing frame-or object-level visual representations.In this paper, we reconsider the multi-modal alignment in VideoQA from feature and sample perspectives to achieve better performance.From the view of feature, we break down the video into trajectories and first leverage trajectory feature in VideoQA to enhance the alignment between two modalities.Moreover, we adopt a heterogeneous graph architecture and design a hierarchical framework to align both trajectory-level and frame-level visual feature with language feature.In addition, we found that VideoQA models are largely dependent on language priors and always neglect visuallanguage interactions.Thus, two effective yet portable training augmentation strategies are designed to strengthen the cross-modal correspondence ability of our model from the view of sample.Extensive results show that our method outperforms all state-of-the-art models on the challenging NExT-QA benchmark.* Long Chen is the corresponding author.… objects trajectories Q: Why did the boy in orange hold a ball on his head?A: want to throw the ball.
Shaoning Xiao, Long Chen 0016, Kaifeng Gao, Yi Yang 0001, Jun Xiao 0001
EMNLP7
2022 Dynamic Feature Pyramid Networks for Detection
abstract
Feature Pyramid Network (FPN) has been a generic feature extractor in computer vision tasks, which utilizes multi-level features to generate discriminative pyramidal representations. However, the way simply using Sum or Concatenate operation on features to integrate multi-scale information is not sufficient to obtain discriminative semantic representations. In this paper, we propose a dynamic feature pyramid network (DyFPN) to merge multi-scale information in both features and weights. DyFPN uses both high-level context features and low-level spatial structural features to obtain dynamic convolution kernel that contains multi-scale information. In this manner, each resolution in the pyramid performs unique and adaptive convolution directly, meanwhile strengthening the information flow. Specially, DyFPN can be regarded as a complementary enhancement to existing feature pyramid networks. We analyze the effective receptive field and attention map of DyFPN. It proves that our method contains more local information and global information compared with merging multi-scale information only on feature level. Benefit from multi-ways of integrating multi-scale information, our method outperforms other existing feature pyramid methods on COCO detection tasks by a large margin.
Kai Zhang 0055, Zheyang Li, Haoji Hu, Bin Li 0025, Wenming Tan, Haixian Lu, Jun Xiao 0001, Ye Ren, Shiliang Pu
ICME7
2022 Deconfounded Value Decomposition for Multi-Agent Reinforcement Learning
abstract
Value decomposition (VD) methods have been widely used in cooperative multi-agent reinforcement learning (MARL), where credit assignment plays an important role in guiding the agents’ decentralized execution. In this paper, we investigate VD from a novel perspective of causal inference. We first show that the environment in existing VD methods is an unobserved confounder as the common cause factor of the global state and the joint value function, which leads to the confounding bias on learning credit assignment. We then present our approach, deconfounded value decomposition (DVD), which cuts off the backdoor confounding path from the global state to the joint value function. The cut is implemented by introducing the trajectory graph, which depends only on the local trajectories, as a proxy confounder. DVD is general enough to be applied to various VD methods, and extensive experiments show that DVD can consistently achieve significant performance gains over different state-of-the-art VD methods on StarCraft II and MACO benchmarks.
Jiahui Li 0003, Kun Kuang 0001, Baoxiang Wang 0001, Furui Liu, Long Chen 0016, Changjie Fan, Fei Wu 0001, Jun Xiao 0001
ICML8
2022 Integrating Object-aware and Interaction-aware Knowledge for Weakly Supervised Scene Graph Generation
abstract
Recently, increasing efforts have been focused on Weakly Supervised Scene Graph Generation (WSSGG). The mainstream solution for WSSGG typically follows the same pipeline: they first align text entities in the weak image-level supervisions (e.g., unlocalized relation triplets or captions) with image regions, and then train SGG models in a fully-supervised manner with aligned instance-level "pseudo" labels. However, we argue that most existing WSSGG works only focus on object-consistency, which means the grounded regions should have the same object category label as text entities. While they neglect another basic requirement for an ideal alignment: interaction-consistency, which means the grounded region pairs should have the same interactions (i.e., visual relations) as text entity pairs. Hence, in this paper, we propose to enhance a simple grounding module with both object-aware and interaction-aware knowledge to acquire more reliable pseudo labels. To better leverage these two types of knowledge, we regard them as two teachers and fuse their generated targets to guide the training process of our grounding module. Specifically, we design two different strategies to adaptively assign weights to different teachers by assessing their reliability on each training sample. Extensive experiments have demonstrated that our method consistently improves WSSGG performance on various kinds of weak supervision.
Long Chen 0016, Yi Yang 0001, Jun Xiao 0001
ACM Multimedia5
2022 Bidirectional Self-Training with Multiple Anisotropic Prototypes for Domain Adaptive Semantic Segmentation
abstract
A thriving trend for domain adaptive segmentation endeavors to generate the high-quality pseudo labels for target domain and retrain the segmentor on them. Under this self-training paradigm, some competitive methods have sought to the latent-space information, which establishes the feature centroids (a.k.a prototypes) of the semantic classes and determines the pseudo label candidates by their distances from these centroids. In this paper, we argue that the latent space contains more information to be exploited thus taking one step further to capitalize on it. Firstly, instead of merely using the source-domain prototypes to determine the target pseudo labels as most of the traditional methods do, we bidirectionally produce the target-domain prototypes to degrade those source features which might be too hard or disturbed for the adaptation. Secondly, existing attempts simply model each category as a single and isotropic prototype while ignoring the variance of the feature distribution, which could lead to the confusion of similar categories. To cope with this issue, we propose to represent each category with multiple and anisotropic prototypes via Gaussian Mixture Model, in order to fit the de facto distribution of source domain and estimate the likelihood of target samples based on the probability density. We apply our method on GTA5->Cityscapes and Synthia->Cityscapes tasks and achieve 61.2% and 62.8% respectively in terms of mean IoU, substantially outperforming other competitive self-training methods. Noticeably, in some categories which severely suffer from the categorical confusion such as "truck" and "bus", our method achieves 56.4% and 68.8% respectively, which further demonstrates the effectiveness of our design. The code and model are available at https://github.com/luyvlei/BiSMAPs.
Yulei Lu, Yawei Luo, Zheyang Li, Yi Yang 0001, Jun Xiao 0001
ACM Multimedia6
2022 Rethinking the Reference-based Distinctive Image Captioning
abstract
Distinctive Image Captioning (DIC) --- generating distinctive captions that describe the unique details of a target image --- has received considerable attention over the last few years. A recent DIC work proposes to generate distinctive captions by comparing the target image with a set of semantic-similar reference images, i.e., reference-based DIC (Ref-DIC). It aims to make the generated captions can tell apart the target and reference images. Unfortunately, reference images used by existing Ref-DIC works are easy to distinguish: these reference images only resemble the target image at scene-level and have few common objects, such that a Ref-DIC model can trivially generate distinctive captions even without considering the reference images. For example, if the target image contains objects "towel'' and "toilet'' while all reference images are without them, then a simple caption "A bathroom with a towel and a toilet'' is distinctive enough to tell apart target and reference images. To ensure Ref-DIC models really perceive the unique objects (or attributes) in target images, we first propose two new Ref-DIC benchmarks. Specifically, we design a two-stage matching mechanism, which strictly controls the similarity between the target and reference images at object-/attribute- level (vs. scene-level). Secondly, to generate distinctive captions, we develop a strong Transformer-based Ref-DIC baseline, dubbed as TransDIC. It not only extracts visual features from the target image, but also encodes the differences between objects in the target and reference images. Finally, for more trustworthy benchmarking, we propose a new evaluation metric named DisCIDEr for Ref-DIC, which evaluates both the accuracy and distinctiveness of the generated captions. Experimental results demonstrate that our TransDIC can generate distinctive captions. Besides, it outperforms several state-of-the-art models on the two new benchmarks over different metrics.
Yangjun Mao, Long Chen 0016, Zhihong Jiang, Jian Shao 0001, Jun Xiao 0001
ACM Multimedia7
2022 Learning Hybrid Behavior Patterns for Multimedia Recommendation
abstract
Multimedia recommendation aims to predict user preferences where users interact with multimodal items. Collaborative filtering based on graph convolutional networks manifests impressive performance gains in multimedia recommendation. This is attributed to the capability of learning good user and item embeddings by aggregating the collaborative signals from high-order neighbors. However, previous researches [37,38] fail to explicitly mine different behavior patterns (i.e., item categories, common user interests) by exploiting user-item and item-item graphs simultaneously, which plays an important role in modeling user preferences. And it is the lack of different behavior pattern constraints and multimodal feature reconciliations that results in performance degradation. Towards this end, We propose a Hybrid Clustering Graph Convolutional Network (HCGCN) for multimedia recommendation. We perform high-order graph convolutions inside user-item clusters and item-item clusters to capture various user behavior patterns. Meanwhile, we design corresponding clustering losses to enhance user-item preference feedback and multimodal representation learning constraint to adjust the modality importance, making more accurate recommendations. Experimental results on three real-world multimedia datasets not only demonstrate the significant improvement of our model over the state-of-the-art methods, but also validate the effectiveness of integrating hybrid user behavior patterns for multimedia recommendation.
Zongshen Mu, Yueting Zhuang, Jun Xiao 0001, Siliang Tang
ACM Multimedia4
2022 Active Learning for Point Cloud Semantic Segmentation via Spatial-Structural Diversity Reasoning
abstract
The expensive annotation cost is notoriously known as the main constraint for the development of the point cloud semantic segmentation technique. Active learning methods endeavor to reduce such cost by selecting and labeling only a subset of the point clouds, yet previous attempts ignore the spatial-structural diversity of the selected samples, inducing the model to select clustered candidates with similar shapes in a local area while missing other representative ones in the global environment. In this paper, we propose a new 3D region-based active learning method to tackle this problem. Dubbed SSDR-AL, our method groups the original point clouds into superpoints and incrementally selects the most informative and representative ones for label acquisition. We achieve the selection mechanism via a graph reasoning network that considers both the spatial and structural diversities of superpoints. To deploy SSDR-AL in a more practical scenario, we design a noise-aware iterative labeling strategy to confront the "noisy annotation'' problem introduced by the previous "dominant labeling'' strategy in superpoints. Extensive experiments on two point cloud benchmarks demonstrate the effectiveness of SSDR-AL in the semantic segmentation task. Particularly, SSDR-AL significantly outperforms the baseline method and reduces the annotation cost by up to $63.0%$ and $24.0%$ when achieving $90%$ performance of fully supervised learning, respectively. Code is available at https://github.com/shaofeifei11/SSDR-AL.
Feifei Shao, Yawei Luo, Ping Liu 0004, Yi Yang 0001, Yulei Lu, Jun Xiao 0001
ACM Multimedia7
2022 Unified Normalization for Accelerating and Stabilizing Transformers
abstract
Solid results from Transformers have made them prevailing architectures in various natural language and vision tasks. As a default component in Transformers, Layer Normalization (LN) normalizes activations within each token to boost the robustness. However, LN requires on-the-fly statistics calculation in inference as well as division and square root operations, leading to inefficiency on hardware. What is more, replacing LN with other hardware-efficient normalization schemes (e.g., Batch Normalization) results in inferior performance, even collapse in training. We find that this dilemma is caused by abnormal behaviors of activation statistics, including large fluctuations over iterations and extreme outliers across layers. To tackle these issues, we propose Unified Normalization (UN), which can speed up the inference by being fused with other linear operations and achieve comparable performance on par with LN. UN strives to boost performance by calibrating the activation and gradient statistics with a tailored fluctuation smoothing strategy. Meanwhile, an adaptive outlier filtration strategy is applied to avoid collapse in training whose effectiveness is theoretically proved and experimentally verified in this paper. We demonstrate that UN can be an efficient drop-in alternative to LN by conducting extensive experiments on language and vision tasks. Besides, we evaluate the efficiency of our method on GPU. Transformers equipped with UN enjoy about 31% inference speedup and nearly 18% memory reduction. Code will be released at https://github.com/hikvision-research/Unified-Normalization.
Kai Zhang 0055, Chaoxiang Lan, Zheyang Li, Wenming Tan, Jun Xiao 0001, Shiliang Pu
ACM Multimedia7
2022 SAViT: Structure-Aware Vision Transformer Pruning via Collaborative Optimization
abstract
Vision Transformers (ViTs) yield impressive performance across various vision tasks. However, heavy computation and memory footprint make them inaccessible for edge devices. Previous works apply importance criteria determined independently by each individual component to prune ViTs. Considering that heterogeneous components in ViTs play distinct roles, these approaches lead to suboptimal performance. In this paper, we introduce joint importance, which integrates essential structural-aware interactions between components for the first time, to perform collaborative pruning. Based on the theoretical analysis, we construct a Taylor-based approximation to evaluate the joint importance. This guides pruning toward a more balanced reduction across all components. To further reduce the algorithm complexity, we incorporate the interactions into the optimization function under some mild assumptions. Moreover, the proposed method can be seamlessly applied to various tasks including object detection. Extensive experiments demonstrate the effectiveness of our method. Notably, the proposed approach outperforms the existing state-of-the-art approaches on ImageNet, increasing accuracy by 0.7% over the DeiT-Base baseline while saving 50% FLOPs. On COCO, we are the first to show that 70% FLOPs of FasterRCNN with ViT backbone can be removed with only 0.3% mAP drop. The code is available at https://github.com/hikvision-research/SAViT.
Chuanyang Zheng, Zheyang Li, Kai Zhang 0055, Wenming Tan, Jun Xiao 0001, Ye Ren, Shiliang Pu
NeurIPS6
2022 Deep Learning for Weakly-Supervised Object Detection and Localization: A Survey
Feifei Shao, Long Chen 0016, Jian Shao 0001, Wei Ji 0008, Shaoning Xiao, Lu Ye, Yueting Zhuang, Jun Xiao 0001
Neurocomputing8
2022 ROBY: Evaluating the adversarial robustness of a deep model by its decision boundaries
Haibo Jin, Jinyin Chen, Haibin Zheng, Zhen Wang 0004, Jun Xiao 0001, Shanqing Yu, Zhaoyan Ming
Inf. Sci.5
2022 Shuhai: A Tool for Benchmarking High Bandwidth Memory on FPGAs
abstract
FPGAs are starting to incorporate High Bandwidth Memory (HBM) to both reduce the memory bandwidth bottleneck encountered in some applications and to provide more capacity to store application state. However, the overall performance characteristics of HBMs are still not well understood, especially in the context of FPGAs, making it difficult to optimize designs relying on HBM. In this article, we bridge the gap between nominal specifications and actual performance by characterizing HBM on a state-of-the-art FPGA, i.e., a Xilinx Alveo U280 featuring a two-stack HBM subsystem. To this end, we have developed Shuhai, a benchmarking tool that throws light on all the subtle details of the performance and usage of HBMs on an FPGA. FPGA-based benchmarking should also provide a more accurate picture of HBM than measuring performance on CPUs/GPUs, since CPUs/GPUs are noisier systems due to their complex control logic and cache hierarchy. Since the memory itself is complex, leveraging custom hardware logic to benchmark it directly from an FPGA provides more details as well as more accurate and deterministic measurements. We observe that 1) HBM is able to provide up to 425 GB/s memory bandwidth, and 2) how HBM is used has a significant impact on the achievable throughput, which in turn demonstrates the importance of unveiling the performance characteristics of HBM so as to use HBM in the right manner. To demonstrate the generality of Shuhai, we also show results for other types of memory, e.g., DDR4, and DDR3, and quantitatively compare the performance characteristics of HBM with those of DDR4 and DDR3.
Hongjing Huang, Zeke Wang, Jie Zhang 0081, Zhenhao He, Chao Wu 0001, Jun Xiao 0001, Gustavo Alonso
IEEE Trans. Computers6
2022 TICS: text-image-based semantic CAPTCHA synthesis via multi-condition adversarial learning
Xinkang Jia, Jun Xiao 0001, Chao Wu 0001
Vis. Comput.2
2021 Ref-NMS: Breaking Proposal Bottlenecks in Two-Stage Referring Expression Grounding
abstract
The prevailing framework for solving referring expression grounding is based on a two-stage process: 1) detecting proposals with an object detector and 2) grounding the referent to one of the proposals. Existing two-stage solutions mostly focus on the grounding step, which aims to align the expressions with the proposals. In this paper, we argue that these methods overlook an obvious mismatch between the roles of proposals in the two stages: they generate proposals solely based on the detection confidence (i.e., expression-agnostic), hoping that the proposals contain all right instances in the expression (i.e., expression-aware). Due to this mismatch, current two-stage methods suffer from a severe performance drop between detected and ground-truth proposals. To this end, we propose Ref-NMS, which is the first method to yield expression-aware proposals at the first stage. Ref-NMS regards all nouns in the expression as critical objects, and introduces a lightweight module to predict a score for aligning each box with a critical object. These scores can guide the NMS operation to filter out the boxes irrelevant to the expression, increasing the recall of critical objects, resulting in a significantly improved grounding performance. Since Ref- NMS is agnostic to the grounding step, it can be easily integrated into any state-of-the-art two-stage method. Extensive ablation studies on several backbones, benchmarks, and tasks consistently demonstrate the superiority of Ref-NMS. Codes are available at: https://github.com/ChopinSharp/ref-nms.
Long Chen 0016, Jun Xiao 0001, Hanwang Zhang, Shih-Fu Chang
AAAI3
2021 Boundary Proposal Network for Two-stage Natural Language Video Localization
abstract
We aim to address the problem of Natural Language Video Localization (NLVL) — localizing the video segment corresponding to a natural language description in a long and untrimmed video. State-of-the-art NLVL methods are almost in one-stage fashion, which can be typically grouped into two categories: 1) anchor-based approach: it first pre-defines a series of video segment candidates (e.g., by sliding window), and then does classification for each candidate; 2) anchor-free approach: it directly predicts the probabilities for each video frame as a boundary or intermediate frame inside the positive segment. However, both kinds of one-stage approaches have inherent drawbacks: the anchor-based approach is susceptible to the heuristic rules, further limiting the capability of handling videos with variant length. While the anchor-free approach fails to exploit the segment-level interaction thus achieving inferior results. In this paper, we propose a novel Boundary Proposal Network (BPNet), a universal two-stage framework that gets rid of the issues mentioned above. Specifically, in the first stage, BPNet utilizes an anchor-free model to generate a group of high-quality candidate video segments with their boundaries. In the second stage, a visual-language fusion layer is proposed to jointly model the multi-modal interaction between the candidate and the language query, followed by a matching score rating layer that outputs the alignment score for each candidate. We evaluate our BPNet on three challenging NLVL benchmarks (i.e., Charades-STA, TACoS and ActivityNet-Captions). Extensive experiments and ablative studies on these datasets demonstrate that the BPNet outperforms the state-of-the-art methods.
Shaoning Xiao, Long Chen 0016, Songyang Zhang 0004, Wei Ji 0008, Jian Shao 0001, Lu Ye, Jun Xiao 0001
AAAI7
2021 Consensus Graph Representation Learning for Better Grounded Image Captioning
abstract
The contemporary visual captioning models frequently hallucinate objects that are not actually in a scene, due to the visual misclassification or over-reliance on priors that resulting in the semantic inconsistency between the visual information and the target lexical words. The most common way is to encourage the captioning model to dynamically link generated object words or phrases to appropriate regions of the image, i.e., the grounded image captioning (GIC). However, GIC utilizes an auxiliary task (grounding objects) that has not solved the key issue of object hallucination, i.e., the semantic inconsistency. In this paper, we take a novel perspective on the issue above: exploiting the semantic coherency between the visual and language modalities. Specifically, we propose the Consensus Rraph Representation Learning framework (CGRL) for GIC that incorporates a consensus representation into the grounded captioning pipeline. The consensus is learned by aligning the visual graph (e.g., scene graph) to the language graph that consider both the nodes and edges in a graph. With the aligned consensus, the captioning model can capture both the correct linguistic characteristics and visual relevance, and then grounding appropriate image regions further. We validate the effectiveness of our model, with a significant decline in object hallucination (-9% CHAIRi) on the Flickr30k Entities dataset. Besides, our CGRL also evaluated by several automatic metrics and human evaluation, the results indicate that the proposed approach can simultaneously improve the performance of image captioning (+2.9 Cider) and grounding (+2.3 F1LOC}).
Wenqiao Zhang, Siliang Tang, Jun Xiao 0001, Yueting Zhuang
AAAI4
2021 Human-Like Controllable Image Captioning With Verb-Specific Semantic Roles
abstract
Controllable Image Captioning (CIC) — generating image descriptions following designated control signals — has received unprecedented attention over the last few years. To emulate the human ability in controlling caption generation, current CIC studies focus exclusively on control signals concerning objective properties, such as contents of interest or descriptive patterns. However, we argue that almost all existing objective control signals have overlooked two indispensable characteristics of an ideal control signal: 1) Event-compatible: all visual contents referred to in a single sentence should be compatible with the described activity. 2) Sample-suitable: the control signals should be suitable for a specific image sample. To this end, we propose a new control signal for CIC: Verb-specific Semantic Roles (VSR). VSR consists of a verb and some semantic roles, which represents a targeted activity and the roles of entities involved in this activity. Given a designated VSR, we first train a grounded semantic role labeling (GSRL) model to identify and ground all entities for each role. Then, we propose a semantic structure planner (SSP) to learn human-like descriptive semantic structures. Lastly, we use a role-shift captioning model to generate the captions. Extensive experiments and ablations demonstrate that our framework can achieve better controllability than several strong base-lines on two challenging CIC benchmarks. Besides, we can generate multi-level diverse captions easily. The code is available at: https://github.com/mad-red/VSR-guided-CIC.
Long Chen 0016, Zhihong Jiang, Jun Xiao 0001, Wei Liu 0005
CVPR3
2021 Natural Language Video Localization with Learnable Moment Proposals
abstract
Given an untrimmed video and a natural language query, Natural Language Video Localization (NLVL) aims to identify the video moment described by the query.To address this task, existing methods can be roughly grouped into two groups: 1) propose-and-rank models first define a set of hand-designed moment candidates and then find out the best-matching one.2) proposal-free models directly predict two temporal boundaries of the referential moment from frames.Currently, almost all the propose-and-rank methods have inferior performance than proposal-free counterparts.In this paper, we argue that propose-and-rank approach is underestimated due to the predefined manners: 1) Hand-designed rules are hard to guarantee the complete coverage of targeted segments.2) Densely sampled candidate moments cause redundant computation and degrade the performance of ranking process.To this end, we propose a novel model termed LP-Net (Learnable Proposal Network for NLVL) with a fixed set of learnable moment proposals.The position and length of these proposals are dynamically adjusted during training process.Moreover, a boundary-aware loss has been proposed to leverage frame-level information and further improve the performance.Extensive ablations on two challenging NLVL benchmarks have demonstrated the effectiveness of LPNet over existing state-of-the-art methods 1 .
Shaoning Xiao, Long Chen 0016, Jian Shao 0001, Yueting Zhuang, Jun Xiao 0001
EMNLP (1)5
2021 Shapley Counterfactual Credits for Multi-Agent Reinforcement Learning
abstract
Centralized Training with Decentralized Execution (CTDE) has been a popular paradigm in cooperative Multi-Agent Reinforcement Learning (MARL) settings and is widely used in many real applications. One of the major challenges in the training process is credit assignment, which aims to deduce the contributions of each agent according to the global rewards. Existing credit assignment methods focus on either decomposing the joint value function into individual value functions or measuring the impact of local observations and actions on the global value function. These approaches lack a thorough consideration of the complicated interactions among multiple agents, leading to an unsuitable assignment of credit and subsequently mediocre results on MARL. We propose Shapley Counterfactual Credit Assignment, a novel method for explicit credit assignment which accounts for the coalition of agents. Specifically, Shapley Value and its desired properties are leveraged in deep MARL to credit any combinations of agents, which grants us the capability to estimate the individual credit for each agent. Despite this capability, the main technical difficulty lies in the computational complexity of Shapley Value who grows factorially as the number of agents. We instead utilize an approximation method via Monte Carlo sampling, which reduces the sample complexity while maintaining its effectiveness. We evaluate our method on StarCraft II benchmarks across different scenarios. Our method outperforms existing cooperative MARL algorithms significantly and achieves the state-of-the-art, with especially large margins on tasks with more severe difficulties.
Jiahui Li 0003, Kun Kuang 0001, Baoxiang Wang 0001, Furui Liu, Long Chen 0016, Fei Wu 0001, Jun Xiao 0001
KDD7
2021 Video Relation Detection via Tracklet based Visual Transformer
abstract
Video Visual Relation Detection (VidVRD), has received significant attention of our community over recent years. In this paper, we apply the state-of-the-art video object tracklet detection pipeline MEGA[7] and deepSORT [27] to generate tracklet proposals. Then we perform VidVRD in a tracklet-based manner without any pre-cutting operations. Specifically, we design a tracklet-based visual Transformer. It contains a temporal-aware decoder which performs feature interactions between the tracklets and learnable predicate query embeddings, and finally predicts the relations. Experimental results strongly demonstrate the superiority of our method, which outperforms other methods by a large margin on the Video Relation Understanding (VRU) Grand Challenge in ACM Multimedia 2021. Codes are released at https://github.com/Dawn-LX/VidVRD-tracklets.
Kaifeng Gao, Long Chen 0016, Jun Xiao 0001
ACM Multimedia4
2021 Instance-wise or Class-wise? A Tale of Neighbor Shapley for Concept-based Explanation
abstract
Interpreting model knowledge is an essential topic to improve human understanding of deep black-box models. Traditional methods contribute to providing intuitive instance-wise explanations which allocating importance scores for low-level features (e.g, pixels for images). To adapt to the human way of thinking, one strand of recent researches has shifted its spotlight to mining important concepts. However, these concept-based interpretation methods focus on computing the contribution of each discovered concept on the class level and can not precisely give instance-wise explanations. Besides, they consider each concept as an independent unit, and ignore the interactions among concepts. To this end, in this paper, we propose a novel COncept-based NEighbor Shapley approach (dubbed as CONE-SHAP) to evaluate the importance of each concept by considering its physical and semantic neighbors, and interpret model knowledge with both instance-wise and class-wise explanations. Thanks to this design, the interactions among concepts in the same image are fully considered. Meanwhile, the computational complexity of Shapley Value is reduced from exponential to polynomial. Moreover, for a more comprehensive evaluation, we further propose three criteria to quantify the rationality of the allocated contributions for the concepts, including coherency, complexity, and faithfulness. Extensive experiments and ablations have demonstrated that our CONE-SHAP algorithm outperforms existing concept-based methods and simultaneously provides precise explanations for each instance and class.
Jiahui Li 0003, Kun Kuang 0001, Lin Li 0065, Long Chen 0016, Songyang Zhang 0004, Jian Shao 0001, Jun Xiao 0001
ACM Multimedia7
2021 Improving Weakly Supervised Object Localization via Causal Intervention
abstract
The recently emerged weakly-supervised object localization (WSOL) methods can learn to localize an object in the image only using image-level labels. Previous works endeavor to perceive the interval objects from the small and sparse discriminative attention map, yet ignoring the co-occurrence confounder (e.g., duck and water), which makes the model inspection (e.g., CAM) hard to distinguish between the object and context. In this paper, we make an early attempt to tackle this challenge via causal intervention (CI). Our proposed method, dubbed CI-CAM, explores the causalities among image features, contexts, and categories to eliminate the biased object-context entanglement in the class activation maps thus improving the accuracy of object localization. Extensive experiments on several benchmarks demonstrate the effectiveness of CI-CAM in learning the clear object boundary from confounding contexts. Particularly, on the CUB-200-2011 which severely suffers from the co-occurrence confounder, CI-CAM significantly outperforms the traditional CAM-based baseline (58.39% vs 52.4% in Top-1 localization accuracy). While in more general scenarios such as ILSVRC 2016, CI-CAM can also perform on par with the state of the arts.
Feifei Shao, Yawei Luo, Lu Ye, Siliang Tang, Yi Yang 0001, Jun Xiao 0001
ACM Multimedia7
2021 Tell and guess: cooperative learning for natural image caption generation with hierarchical refined attention
Wenqiao Zhang, Siliang Tang, Jiajie Su, Jun Xiao 0001, Yueting Zhuang
Multim. Tools Appl.4
2021 Explore Video Clip Order With Self-Supervised and Curriculum Learning for Video Applications
abstract
We present a self-supervised spatiotemporal learning approach by exploring the temporal coherence of videos. The chronological order of shuffled clips from the video is used as the supervisory signal to guide the 3D Convolutional Neural Networks (CNNs) to learn meaningful visual knowledge. Unlike the existing approaches which use frames, we utilize dynamic video clips to reduce the uncertainty of order. We test three types of representative 3D CNNs, all of which benefit from the proposed approach. The learned 3D CNNs can be used either as a feature extractor or a pre-trained model for further fine-tuning on downstream tasks. We also propose two curriculum learning strategies to make the 3D CNNs easier to train and get the state-of-the-art results in nearest neighbor retrieval and action recognition tasks compared with other self-supervised learning methods. Meanwhile, it is further extended to the field of visual question answering application and has achieved promising results. Besides, comprehensive and extensive experimental results and analyses are provided for readers to better understand the video clip order we explore with self-supervised and curriculum learning for video application.
Jun Xiao 0001, Lin Li 0065, Dejing Xu, Chengjiang Long, Jian Shao 0001, Shiliang Pu, Yueting Zhuang
IEEE Trans. Multim.1
2020 Rethinking the Bottom-Up Framework for Query-Based Video Localization
abstract
In this paper, we focus on the task query-based video localization, i.e., localizing a query in a long and untrimmed video. The prevailing solutions for this problem can be grouped into two categories: i) Top-down approach: It pre-cuts the video into a set of moment candidates, then it does classification and regression for each candidate; ii) Bottom-up approach: It injects the whole query content into each video frame, then it predicts the probabilities of each frame as a ground truth segment boundary (i.e., start or end). Both two frameworks have respective shortcomings: the top-down models suffer from heavy computations and they are sensitive to the heuristic rules, while the performance of bottom-up models is behind the performance of top-down counterpart thus far. However, we argue that the performance of bottom-up framework is severely underestimated by current unreasonable designs, including both the backbone and head network. To this end, we design a novel bottom-up model: Graph-FPN with Dense Predictions (GDP). For the backbone, GDP firstly generates a frame feature pyramid to capture multi-level semantics, then it utilizes graph convolution to encode the plentiful scene relationships, which incidentally mitigates the semantic gaps in the multi-scale feature pyramid. For the head network, GDP regards all frames falling in the ground truth segment as the foreground, and each foreground frame regresses the unique distances from its location to bi-directional boundaries. Extensive experiments on two challenging query-based video localization tasks (natural language video localization and video relocalization), involving four challenging benchmarks (TACoS, Charades-STA, ActivityNet Captions, and Activity-VRL), have shown that GDP surpasses the state-of-the-art top-down models.
Long Chen 0016, Chujie Lu, Siliang Tang, Jun Xiao 0001, Chilie Tan
AAAI4
2020 Counterfactual Samples Synthesizing for Robust Visual Question Answering
abstract
Despite Visual Question Answering (VQA) has realized impressive progress over the last few years, today's VQA models tend to capture superficial linguistic correlations in the train set and fail to generalize to the test set with different QA distributions. To reduce the language biases, several recent works introduce an auxiliary question-only model to regularize the training of targeted VQA model, and achieve dominating performance on VQA-CP. However, since the complexity of design, current methods are unable to equip the ensemble-based models with two indispensable characteristics of an ideal VQA model: 1) visual-explainable: the model should rely on the right visual regions when making decisions. 2) question-sensitive: the model should be sensitive to the linguistic variations in question. To this end, we propose a model-agnostic Counterfactual Samples Synthesizing (CSS) training scheme. The CSS generates numerous counterfactual training samples by masking critical objects in images or words in questions, and assigning different ground-truth answers. After training with the complementary samples (ie, the original and generated samples), the VQA models are forced to focus on all critical objects and words, which significantly improves both visual-explainable and question-sensitive abilities. In return, the performance of these models is further boosted. Extensive ablations have shown the effectiveness of CSS. Particularly, by building on top of the model LMH, we achieve a record-breaking performance of 58.95% on VQA-CP v2, with 6.5% gains.
Long Chen 0016, Jun Xiao 0001, Hanwang Zhang, Shiliang Pu, Yueting Zhuang
CVPR3
2020 De-Biased Court's View Generation with Causality
abstract
Yiquan Wu, Kun Kuang, Yating Zhang, Xiaozhong Liu, Changlong Sun, Jun Xiao, Yueting Zhuang, Luo Si, Fei Wu. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Yiquan Wu 0001, Kun Kuang 0001, Xiaozhong Liu 0001, Changlong Sun, Jun Xiao 0001, Yueting Zhuang, Luo Si, Fei Wu 0001
EMNLP (1)6
2020 Hierarchical Attention Based Spatial-Temporal Graph-to-Sequence Learning for Grounded Video Description
abstract
The task of Grounded Video Description~(GVD) is to generate sentences whose objects can be grounded with the bounding boxes in the video frames. Existing works often fail to exploit structural information both in modeling the relationships among the region proposals and in attending them for text generation. To address these issues, we cast the GVD task as a spatial-temporal Graph-to-Sequence learning problem, where we model video frames as spatial-temporal sequence graph in order to better capture implicit structural relationships. In particular, we exploit two ways to construct a sequence graph that captures spatial-temporal correlations among different objects in each frame and further present a novel graph topology refinement technique to discover optimal underlying graph structure. In addition, we also present hierarchical attention mechanism to attend sequence graph in different resolution levels for better generating the sentences. Our extensive experiments demonstrate the effectiveness of our proposed method compared to state-of-the-art methods.
Lingfei Wu 0001, Fangli Xu, Siliang Tang, Jun Xiao 0001, Yueting Zhuang
IJCAI5
2020 Topic Adaptation and Prototype Encoding for Few-Shot Visual Storytelling
abstract
Visual Storytelling~(VIST) is a task to tell a narrative story about a certain topic according to the given photo stream. The existing studies focus on designing complex models, which rely on a huge amount of human-annotated data. However, the annotation of VIST is extremely costly and many topics cannot be covered in the training dataset due to the long-tail topic distribution. In this paper, we focus on enhancing the generalization ability of the VIST model by considering the few-shot setting. Inspired by the way humans tell a story, we propose a topic adaptive storyteller to model the ability of inter-topic generalization. In practice, we apply the gradient-based meta-learning algorithm on multi-modal seq2seq models to endow the model the ability to adapt quickly from topic to topic. Besides, We further propose a prototype encoding structure to model the ability of intra-topic derivation. Specifically, we encode and restore the few training story text to serve as a reference to guide the generation at inference time. Experimental results show that topic adaptation and prototype encoding structure mutually bring benefit to the few-shot model on BLEU and METEOR metric. The further case study shows that the stories generated after few-shot adaptation are more relative and expressive.
Jiacheng Li 0002, Siliang Tang, Juncheng Li 0006, Jun Xiao 0001, Fei Wu 0001, Shiliang Pu, Yueting Zhuang
ACM Multimedia4
2020 Photo Stream Question Answer
abstract
Understanding and reasoning over partially observed visual clues are often regarded as a challenging real-world problem even for human beings. In this paper, we present a new visual question answering (VQA) task -- Photo Stream QA, which aims to answer the open-ended questions about a narrative photo stream. Photo Stream QA is more challenging and interesting than the existing VQA tasks, since the temporal and visual variance among photos in the stream is huge and hard to observe. Therefore, instead of learning simple vision-text mappings, the AI algorithms must fill these variance gaps with more recollection, reasoning, even the knowledge from our daily experiences. To tackle the problems in Photo Stream QA, we propose an end-to-end baseline (E-TAA) with a novel Experienced Unit (E-unit) and Three-stage Alternating Attention (TAA). E-unit yields a better visual representation which captures the temporal semantic relation among visual clues in the photo stream, while TAA creates three levels of attention that gradually refines visual features by using the textual representation from the question as the guidance. Experimental results on our developed dataset demonstrate that, as the first attempt at the Photo Stream QA task, E-TAA provides promising results outperforming all the other baseline methods.
Wenqiao Zhang, Siliang Tang, Yanpeng Cao, Jun Xiao 0001, Shiliang Pu, Fei Wu 0001, Yueting Zhuang
ACM Multimedia4
2020 Relational Graph Learning for Grounded Video Description Generation
abstract
Grounded video description (GVD) encourages captioning models to attend to appropriate video regions (e.g., objects) dynamically and generate a description. Such a setting can help explain the decisions of captioning models and prevents the model from hallucinating object words in its description. However, such design mainly focuses on object word generation and thus may ignore fine-grained information and suffer from missing visual concepts. Moreover, relational words (e.g., 'jump left or right') are usual spatio-temporal inference results, i.e., these words cannot be grounded on certain spatial regions. To tackle the above limitations, we design a novel relational graph learning framework for GVD, in which a language-refined scene graph representation is designed to explore fine-grained visual concepts. Furthermore, the refined graph can be regarded as relational inductive knowledge to assist captioning models in selecting the relevant information it needs to generate correct words. We validate the effectiveness of our model through automatic metrics and human evaluation, and the results indicate that our approach can generate more fine-grained and accurate description, and it solves the problem of object hallucination to some extent.
Wenqiao Zhang, Xin Wang 0061, Siliang Tang, Haizhou Shi, Jun Xiao 0001, Yueting Zhuang, William Yang Wang
ACM Multimedia6
2020 Hierarchical Fashion Graph Network for Personalized Outfit Recommendation
abstract
Fashion outfit recommendation has attracted increasing attentions from online shopping services and fashion communities.Distinct from other scenarios (e.g., social networking or content sharing) which recommend a single item (e.g., a friend or picture) to a user, outfit recommendation predicts user preference on a set of well-matched fashion items. Hence, performing high-quality personalized outfit recommendation should satisfy two requirements -- 1) the nice compatibility of fashion items and 2) the consistence with user preference. However, present works focus mainly on one of the requirements and only consider either user-outfit or outfit-item relationships, thereby easily leading to suboptimal representations and limiting the performance.
Xiang Wang 0010, Xiangnan He 0001, Long Chen 0016, Jun Xiao 0001, Tat-Seng Chua
SIGIR5
2020 Abstractive meeting summarization by hierarchical adaptive segmental network learning with multiple revising steps
Jiyuan Zheng, Zhou Zhao 0001, Zehan Song, Min Yang 0007, Jun Xiao 0001
Neurocomputing5
2020 Video question answering via grounded cross-attention network learning
Yunan Ye, Xufeng Qian, Siliang Tang, Shiliang Pu, Jun Xiao 0001
Inf. Process. Manag.7
2020 Multi-platform data collection for public service with Pay-by-Data
Chao Wu 0001, Simon Hu 0001, Chun-Hsiang Lee, Jun Xiao 0001
Multim. Tools Appl.4
2020 Hierarchical Temporal Fusion of Multi-grained Attention Features for Video Question Answering
Shaoning Xiao, Yunan Ye, Long Chen 0016, Shiliang Pu, Zhou Zhao 0001, Jian Shao 0001, Jun Xiao 0001
Neural Process. Lett.8
2020 Open-Ended Video Question Answering via Multi-Modal Conditional Adversarial Networks
abstract
As a challenging task in visual information retrieval, open-ended long-form video question answering automatically generates the natural language answer from the referenced video content according to the given question. However, the existing video question answering works mainly focus on the short-form video, which may be ineffectively applied for long-form video question answering directly, due to the insufficiency of modeling the semantic representation of long-form video content. In this paper, we study the problem of open-ended long-form video question answering from the viewpoint of hierarchical multimodal conditional adversarial network learning. We propose the hierarchical attentional encoder network to learn the joint representation of long-form video content and given question with adaptive video segmentation. We then devise the reinforced decoder network to generate the natural language answer for openended video question answering with multi-modal conditional adversarial network learning. We construct three large-scale open-ended video question answering datasets. The extensive experiments validate the effectiveness of our method.
Zhou Zhao 0001, Shuwen Xiao, Zehan Song, Chujie Lu, Jun Xiao 0001, Yueting Zhuang
IEEE Trans. Image Process.5
2020 Multichannel Attention Refinement for Video Question Answering
abstract
Video Question Answering (VideoQA) is the extension of image question answering (ImageQA) in the video domain. Methods are required to give the correct answer after analyzing the provided video and question in this task. Comparing to ImageQA, the most distinctive part is the media type. Both tasks require the understanding of visual media, but VideoQA is much more challenging, mainly because of the complexity and diversity of videos. Particularly, working with the video needs to model its inherent temporal structure and analyze the diverse information it contains. In this article, we propose to tackle the task from a multichannel perspective. Appearance, motion, and audio features are extracted from the video, and question-guided attentions are refined to generate the expressive clues that support the correct answer. We also incorporate the relevant text information acquired from Wikipedia as an attempt to extend the capability of the method. Experiments on TGIF-QA and ActivityNet-QA datasets show the advantages of our method compared to existing methods. We also demonstrate the effectiveness and interpretability of our method by analyzing the refined attention weights during the question-answering procedure.
Yueting Zhuang, Dejing Xu, Wenzhuo Cheng, Zhou Zhao 0001, Shiliang Pu, Jun Xiao 0001
ACM Trans. Multim. Comput. Commun. Appl.7
2019 Self-Supervised Spatiotemporal Learning via Video Clip Order Prediction
abstract
We propose a self-supervised spatiotemporal learning technique which leverages the chronological order of videos. Our method can learn the spatiotemporal representation of the video by predicting the order of shuffled clips from the video. The category of the video is not required, which gives our technique the potential to take advantage of infinite unannotated videos. There exist related works which use frames, while compared to frames, clips are more consistent with the video dynamics. Clips can help to reduce the uncertainty of orders and are more appropriate to learn a video representation. The 3D convolutional neural networks are utilized to extract features for clips, and these features are processed to predict the actual order. The learned representations are evaluated via nearest neighbor retrieval experiments. We also use the learned networks as the pre-trained models and finetune them on the action recognition task. Three types of 3D convolutional neural networks are tested in experiments, and we gain large improvements compared to existing self-supervised methods.
Dejing Xu, Jun Xiao 0001, Zhou Zhao 0001, Jian Shao 0001, Di Xie, Yueting Zhuang
CVPR2
2019 Video Dialog via Progressive Inference and Cross-Transformer
abstract
Weike Jin, Zhou Zhao, Mao Gu, Jun Xiao, Furu Wei, Yueting Zhuang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Weike Jin, Zhou Zhao 0001, Mao Gu, Jun Xiao 0001, Furu Wei, Yueting Zhuang
EMNLP/IJCNLP (1)4
2019 DEBUG: A Dense Bottom-Up Grounding Approach for Natural Language Video Localization
abstract
Chujie Lu, Long Chen, Chilie Tan, Xiaolin Li, Jun Xiao. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Chujie Lu, Long Chen 0016, Chilie Tan, Jun Xiao 0001
EMNLP/IJCNLP (1)5
2019 Counterfactual Critic Multi-Agent Training for Scene Graph Generation
abstract
Scene graphs --- objects as nodes and visual relationships as edges --- describe the whereabouts and interactions of objects in an image for comprehensive scene understanding. To generate coherent scene graphs, almost all existing methods exploit the fruitful visual context by modeling message passing among objects. For example, ``person'' on ``bike'' can help to determine the relationship ``ride'', which in turn contributes to the confidence of the two objects. However, we argue that the visual context is not properly learned by using the prevailing cross-entropy based supervised learning paradigm, which is not sensitive to graph inconsistency: errors at the hub or non-hub nodes should not be penalized equally. To this end, we propose a Counterfactual critic Multi-Agent Training (CMAT) approach. CMAT is a multi-agent policy gradient method that frames objects into cooperative agents, and then directly maximizes a graph-level metric as the reward. In particular, to assign the reward properly to each agent, CMAT uses a counterfactual baseline that disentangles the agent-specific reward by fixing the predictions of other agents. Extensive validations on the challenging Visual Genome benchmark show that CMAT achieves a state-of-the-art performance by significant gains under various settings and metrics.
Long Chen 0016, Hanwang Zhang, Jun Xiao 0001, Xiangnan He 0001, Shiliang Pu, Shih-Fu Chang
ICCV3
2019 Weak Supervision Enhanced Generative Network for Question Generation
abstract
Automatic question generation according to an answer within the given passage is useful for many applications, such as question answering system, dialogue system, etc. Current neural-based methods mostly take two steps which extract several important sentences based on the candidate answer through manual rules or supervised neural networks and then use an encoder-decoder framework to generate questions about these sentences. These approaches still acquire two steps and neglect the semantic relations between the answer and the context of the whole passage which is sometimes necessary for answering the question. To address this problem, we propose the Weakly Supervision Enhanced Generative Network (WeGen) which automatically discovers relevant features of the passage given the answer span in a weakly supervised manner to improve the quality of generated questions. More specifically, we devise a discriminator, Relation Guider, to capture the relations between the passage and the associated answer and then the Multi-Interaction mechanism is deployed to transfer the knowledge dynamically for our question generation system. Experiments show the effectiveness of our method in both automatic evaluations and human evaluations.
Jiyuan Zheng, Qijiong Liu, Zhou Zhao 0001, Jun Xiao 0001, Yueting Zhuang
IJCAI5
2019 Multi-interaction Network with Object Relation for Video Question Answering
abstract
Video question answering is an important task for testing machine's ability of video understanding. The existing methods normally focus on the combination of recurrent and convolutional neural networks to capture spatial and temporal information of the video. Recently, some work has also shown that using attention mechanism can achieve better performance. In this paper, we propose a new model called Multi-interaction network for video question answering. There are two types of interactions in our model. The first type is the multi-modal interaction between the visual and textual information. The second type is the multi-level interaction inside the multi-modal interaction. Specifically, instead of using original self-attention, we propose a new attention mechanism called multi-interaction, which can capture both element-wise and segment-wise sequence interactions, simultaneously. And in addition to the normal frame-level interaction, we also take the object relations into consideration, in order to obtain more fine-grained information, such as motions and other potential relations among these objects. We evaluate our method on TGIF-QA and other two video QA datasets. The qualitative and quantitative experimental results show the effectiveness of our model, which achieves the new state-of-the-art performance.
Weike Jin, Zhou Zhao 0001, Mao Gu, Jun Yu 0002, Jun Xiao 0001, Yueting Zhuang
ACM Multimedia5
2019 Video Relation Detection with Spatio-Temporal Graph
abstract
What we perceive from visual content are not only collections of objects but the interactions between them. Visual relations, denoted by the triplet , could convey a wealth of information for visual understanding. Different from static images and because of the additional temporal channel, dynamic relations in videos are often correlated in both spatial and temporal dimensions, which make the relation detection in videos a more complex and challenging task. In this paper, we abstract videos into fully-connected spatial-temporal graphs. We pass message and conduct reasoning in these 3D graphs with a novel VidVRD model using graph convolution network. Our model can take advantage of spatial-temporal contextual cues to make better predictions on objects as well as their dynamic relationships. Furthermore, an online association method with a siamese network is proposed for accurate relation instances association. By combining our model (VRD-GCN) and the proposed association method, our framework for video relation detection achieves the best performance in the latest benchmarks. We validate our approach on benchmark ImageNet-VidVRD dataset. The experimental results show that our framework outperforms the state-of-the-art by a large margin and a series of ablation studies demonstrate our method's effectiveness.
Xufeng Qian, Yueting Zhuang, Shaoning Xiao, Shiliang Pu, Jun Xiao 0001
ACM Multimedia6
2019 Video Dialog via Multi-Grained Convolutional Self-Attention Context Networks
abstract
Video dialog is a new and challenging task, which requires an AI agent to maintain a meaningful dialog with humans in natural language about video contents. Specifically, given a video, a dialog history and a new question about the video, the agent has to combine video information with dialog history to infer the answer. And due to the complexity of video information, the methods of image dialog might be ineffectively applied directly to video dialog. In this paper, we propose a novel approach for video dialog called multi-grained convolutional self-attention context network, which combines video information with dialog history. Instead of using RNN to encode the sequence information, we design a multi-grained convolutional self-attention mechanism to capture both element and segment level interactions which contain multi-grained sequence information. Then, we design a hierarchical dialog history encoder to learn the context-aware question representation and a two-stream video encoder to learn the context-aware video representation. We evaluate our method on two large-scale datasets. Due to the flexibility and parallelism of the new attention mechanism, our method can achieve higher time efficiency, and the extensive experiments also show the effectiveness of our method.
Weike Jin, Zhou Zhao 0001, Mao Gu, Jun Yu 0002, Jun Xiao 0001, Yueting Zhuang
SIGIR5
2019 An artificial intelligence based data-driven approach for design ideation
Liuqing Chen 0002, Pan Wang 0005, Hao Dong 0003, Feng Shi 0007, Yike Guo, Peter R. N. Childs, Jun Xiao 0001, Chao Wu 0001
J. Vis. Commun. Image Represent.8
2019 Adversarial learning for viewpoints invariant 3D human pose estimation
Jun Xiao 0001, Di Xie, Jian Shao 0001
J. Vis. Commun. Image Represent.2
2019 Explorations of skeleton features for LSTM-based action recognition
Jia-geng Feng, Songyang Zhang 0004, Jun Xiao 0001
Multim. Tools Appl.3
2019 Video Question Answering via Knowledge-based Progressive Spatial-Temporal Attention Network
abstract
Visual Question Answering (VQA) is a challenging task that has gained increasing attention from both the computer vision and the natural language processing communities in recent years. Given a question in natural language, a VQA system is designed to automatically generate the answer according to the referenced visual content. Though there recently has been much intereset in this topic, the existing work of visual question answering mainly focuses on a single static image, which is only a small part of the dynamic and sequential visual data in the real world. As a natural extension, video question answering (VideoQA) is less explored. Because of the inherent temporal structure in the video, the approaches of ImageQA may be ineffectively applied to video question answering. In this article, we not only take the spatial and temporal dimension of video content into account but also employ an external knowledge base to improve the answering ability of the network. More specifically, we propose a knowledge-based progressive spatial-temporal attention network to tackle this problem. We obtain both objects and region features of the video frames from a region proposal network. The knowledge representation is generated by a word-level attention mechanism using the comment information of each object that is extracted from DBpedia. Then, we develop a question-knowledge-guided progressive spatial-temporal attention network to learn the joint video representation for video question answering task. We construct a large-scale video question answering dataset. The extensive experiments based on two different datasets validate the effectiveness of our method.
Weike Jin, Zhou Zhao 0001, Jun Xiao 0001, Yueting Zhuang
ACM Trans. Multim. Comput. Commun. Appl.5
2018 Zero-Shot Visual Recognition Using Semantics-Preserving Adversarial Embedding Networks
abstract
We propose a novel framework called Semantics-Preserving Adversarial Embedding Network (SP-AEN) for zero-shot visual recognition (ZSL), where test images and their classes are both unseen during training. SP-AEN aims to tackle the inherent problem - semantic loss - in the prevailing family of embedding-based ZSL, where some semantics would be discarded during training if they are non-discriminative for training classes, but could become critical for recognizing test classes. Specifically, SP-AEN prevents the semantic loss by introducing an independent visual-to-semantic space embedder which disentangles the semantic space into two subspaces for the two arguably conflicting objectives: classification and reconstruction. Through adversarial learning of the two subspaces, SP-AEN can transfer the semantics from the reconstructive subspace to the discriminative one, accomplishing the improved zero-shot recognition of unseen classes. Comparing with prior works, SP-AEN can not only improve classification but also generate photo-realistic images, demonstrating the effectiveness of semantic preservation. On four popular benchmarks: CUB, AWA, SUN and aPY, SP-AEN considerably outperforms other state-of-the-art methods by an absolute performance difference of 12.2%, 9.3%, 4.0% and 3.6% in terms of harmonic mean values [62].
Long Chen 0016, Hanwang Zhang, Jun Xiao 0001, Wei Liu 0005, Shih-Fu Chang
CVPR3
2018 Multi-Turn Video Question Answering via Multi-Stream Hierarchical Attention Context Network
abstract
Conversational video question answering is a challenging task in visual information retrieval, which generates the accurate answer from the referenced video contents according to the visual conversation context and given question. However, the existing visual question answering methods mainly tackle the problem of single-turn video question answering, which may be ineffectively applied for multi-turn video question answering directly, due to the insufficiency of modeling the sequential conversation context. In this paper, we study the problem of multi-turn video question answering from the viewpoint of multi-step hierarchical attention context network learning. We first propose the hierarchical attention context network for context-aware question understanding by modeling the hierarchically sequential conversation context structure. We then develop the multi-stream spatio-temporal attention network for learning the joint representation of the dynamic video contents and context-aware question embedding. We next devise the hierarchical attention context network learning method with multi-step reasoning process for multi-turn video question answering. We construct two large-scale multi-turn video question answering datasets. The extensive experiments show the effectiveness of our method.
Zhou Zhao 0001, Xinghua Jiang, Deng Cai 0001, Jun Xiao 0001, Xiaofei He 0001, Shiliang Pu
IJCAI4
2018 Attentional Image Retweet Modeling via Multi-Faceted Ranking Network Learning
abstract
Retweet prediction is a challenging problem in social media sites (SMS). In this paper, we study the problem of image retweet prediction in social media, which predicts the image sharing behavior that the user reposts the image tweets from their followees. Unlike previous studies, we learn user preference ranking model from their past retweeted image tweets in SMS. We first propose heterogeneous image retweet modeling network (IRM) that exploits users' past retweeted image tweets with associated contexts, their following relations in SMS and preference of their followees. We then develop a novel attentional multi-faceted ranking network learning framework with multi-modal neural networks for the proposed heterogenous IRM network to learn the joint image tweet representations and user preference representations for prediction task. The extensive experiments on a large-scale dataset from Twitter site shows that our method achieves better performance than other state-of-the-art solutions to the problem.
Zhou Zhao 0001, Lingtao Meng, Jun Xiao 0001, Min Yang 0007, Fei Wu 0001, Deng Cai 0001, Xiaofei He 0001, Yueting Zhuang
IJCAI3
2018 Fusing Geometric Features for Skeleton-Based Action Recognition Using Multilayer LSTM Networks
abstract
Recent skeleton-based action recognition approaches achieve great improvement by using recurrent neural network (RNN) models. Currently, these approaches build an end-to-end network from coordinates of joints to class categories and improve accuracy by extending RNN to spatial domains. First, while such well-designed models and optimization strategies explore relations between different parts directly from joint coordinates, we provide a simple universal spatial modeling method perpendicular to the RNN model enhancement. Specifically, according to the evolution of previous work, we select a set of simple geometric features, and then separately feed each type of features to a three-layer LSTM framework. Second, we propose a multistream LSTM architecture with a new smoothed score fusion technique to learn classification from different geometric feature streams. Furthermore, we observe that the geometric relational features based on distances between joints and selected lines outperform other features and the fusion results achieve the state-of-the-art performance on four datasets. We also show the sparsity of input gate weights in the first LSTM layer trained by geometric features and demonstrate that utilizing joint-line distances as input require less data for training.
Songyang Zhang 0004, Yang Yang 0009, Jun Xiao 0001, Xiaoming Liu 0002, Yi Yang 0001, Di Xie, Yueting Zhuang
IEEE Trans. Multim.3
2017 Integrating Side Information for Boosting Machine Comprehension
abstract
Machine Reading and Comprehension recently has drawn a fair amount of attention in the field of natural language processing. In this paper, we consider integrating side information to improve machine comprehension on answering cloze-style questions more precisely. To leverage the external information, we present a novel attention-based architecture which could feed the side information representations into word level embeddings to explore the comprehension performance. Our experiments show consistent improvements of our model over various baselines.
Min Yang 0007, Zhou Zhao 0001, Jun Xiao 0001, Yueting Zhuang
CIKM5
2017 SCA-CNN: Spatial and Channel-Wise Attention in Convolutional Networks for Image Captioning
abstract
Visual attention has been successfully applied in structural prediction tasks such as visual captioning and question answering. Existing visual attention models are generally spatial, i.e., the attention is modeled as spatial probabilities that re-weight the last conv-layer feature map of a CNN encoding an input image. However, we argue that such spatial attention does not necessarily conform to the attention mechanism - a dynamic feature extractor that combines contextual fixations over time, as CNN features are naturally spatial, channel-wise and multi-layer. In this paper, we introduce a novel convolutional neural network dubbed SCA-CNN that incorporates Spatial and Channel-wise Attentions in a CNN. In the task of image captioning, SCA-CNN dynamically modulates the sentence generation context in multi-layer feature maps, encoding where (i.e., attentive spatial locations at multiple layers) and what (i.e., attentive channels) the visual attention is. We evaluate the proposed SCA-CNN architecture on three benchmark image captioning datasets: Flickr8K, Flickr30K, and MSCOCO. It is consistently observed that SCA-CNN significantly outperforms state-of-the-art visual attention-based image captioning methods.
Long Chen 0016, Hanwang Zhang, Jun Xiao 0001, Liqiang Nie, Jian Shao 0001, Wei Liu 0005, Tat-Seng Chua
CVPR3
2017 Graph-theoretic spatiotemporal context modeling for video saliency detection
abstract
As an important and challenging problem in computer vision, video saliency detection is typically cast as a spatiotemporal context modeling problem over consecutive frames. As a result, a key issue in video saliency detection is how to effectively capture the intrinsical properties of atomic video structures as well as their associated contextual interactions along the spatial and temporal dimensions. Motivated by this observation, we propose a graph-theoretic video saliency detection approach based on adaptive video structure discovery, which is carried out within a spatiotemporal atomic graph. Through graph-based manifold propagation, the proposed approach is capable of effectively modeling the semantically contextual interactions among atomic video structures for saliency detection while preserving spatial smoothness and temporal consistency. Experiments demonstrate the effectiveness of the proposed approach over several benchmark datasets.
Lina Wei, Xi Li 0001, Fei Wu 0001, Jun Xiao 0001
ICIP5
2017 Attentional Factorization Machines: Learning the Weight of Feature Interactions via Attention Networks
abstract
Factorization Machines (FMs) are a supervised learning approach that enhances the linear regression model by incorporating the second-order feature interactions. Despite effectiveness, FM can be hindered by its modelling of all feature interactions with the same weight, as not all feature interactions are equally useful and predictive. For example, the interactions with useless features may even introduce noises and adversely degrade the performance. In this work, we improve FM by discriminating the importance of different feature interactions. We propose a novel model named Attentional Factorization Machine (AFM), which learns the importance of each feature interaction from data via a neural attention network. Extensive experiments on two real-world datasets demonstrate the effectiveness of AFM. Empirically, it is shown on regression task AFM betters FM with a 8.6% relative improvement, and consistently outperforms the state-of-the-art deep learning methods Wide&Deep [Cheng et al., 2016] and DeepCross [Shan et al., 2016] with a much simpler structure and fewer model parameters. Our implementation of AFM is publicly available at: https://github.com/hexiangnan/attentional_factorization_machine
Jun Xiao 0001, Xiangnan He 0001, Hanwang Zhang, Fei Wu 0001, Tat-Seng Chua
IJCAI1
2017 Video Question Answering via Gradually Refined Attention over Appearance and Motion
abstract
Recently image question answering (ImageQA) has gained lots of attention in the research community. However, as its natural extension, video question answering (VideoQA) is less explored. Although both tasks look similar, VideoQA is more challenging mainly because of the complexity and diversity of videos. As such, simply extending the ImageQA methods to videos is insufficient and suboptimal. Particularly, working with the video needs to model its inherent temporal structure and analyze the diverse information it contains. In this paper, we consider exploiting the appearance and motion information resided in the video with a novel attention mechanism. More specifically, we propose an end-to-end model which gradually refines its attention over the appearance and motion features of the video using the question as guidance. The question is processed word by word until the model generates the final optimized attention. The weighted representation of the video, as well as other contextual information, are used to generate the answer. Extensive experiments show the advantages of our model compared to other baseline models. We also demonstrate the effectiveness of our model by analyzing the refined attention weights during the question answering procedure.
Dejing Xu, Zhou Zhao 0001, Jun Xiao 0001, Fei Wu 0001, Hanwang Zhang, Xiangnan He 0001, Yueting Zhuang
ACM Multimedia3
2017 ENCORE: External Neural Constraints Regularized Distant Supervision for Relation Extraction
abstract
Distant Supervision is a widely used approach for training relation extraction models. It generates noisy training samples by heuristically labeling a corpus using an existing knowledge base. Previous noise reduction methods for distant supervision fail to utilize information such as data credibility and sample confidence. In this paper, we proposed a novel neural framework, named ENCORE (External Neural COnstraints REgularized distant supervision), which allows an integration of other information for standard DS through regularizations under multiple external neural networks. In ENCORE, a teacher-student co-training mechanism is used to iterative distilling information from external neural networks to an existing relation extraction model. The experiment results demonstrated that without increasing any data or reshaping its original structure, ENCORE enhanced a CNN based relation extraction model for over 12%. The enhanced model also outperforms the state-of-the-art relation extraction method on the same dataset.
Siliang Tang, Jinjian Zhang, Fei Wu 0001, Jun Xiao 0001, Yueting Zhuang
SIGIR5
2017 Video Question Answering via Attribute-Augmented Attention Network Learning
abstract
Video Question Answering is a challenging problem in visual information retrieval, which provides the answer to the referenced video content according to the question. However, the existing visual question answering approaches mainly tackle the problem of static image question, which may be ineffectively for video question answering due to the insufficiency of modeling the temporal dynamics of video contents. In this paper, we study the problem of video question answering by modeling its temporal dynamics with frame-level attention mechanism. We propose the attribute-augmented attention network learning framework that enables the joint frame-level attribute detection and unified video representation learning for video question answering. We then incorporate the multi-step reasoning process for our proposed attention network to further improve the performance. We construct a large-scale video question answering dataset. We conduct the experiments on both multiple-choice and open-ended video question answering tasks to show the effectiveness of the proposed method.
Yunan Ye, Zhou Zhao 0001, Long Chen 0016, Jun Xiao 0001, Yueting Zhuang
SIGIR5
2017 Learning Max-Margin GeoSocial Multimedia Network Representations for Point-of-Interest Suggestion
abstract
With the rapid development of mobile devices, point-of-interest (POI) suggestion has become a popular online web service, which provides attractive and interesting locations to users. In order to provide interesting POIs, many existing POI recommendation works learn the latent representations of users and POIs from users' past visiting POIs, which suffers from the sparsity problem of POI data. In this paper, we consider the problem of POI suggestion from the viewpoint of learning geosocial multimedia network representations. We propose a novel max-margin metric geosocial multimedia network representation learning framework by exploiting users' check-in behavior and their social relations. We then develop a random-walk based learning method with max-margin metric network embedding. We evaluate the performance of our method on a large-scale geosocial multimedia network dataset and show that our method achieves the best performance than other state-of-the-art solutions.
Zhou Zhao 0001, Hanqing Lu, Min Yang 0007, Jun Xiao 0001, Fei Wu 0001, Yueting Zhuang
SIGIR5
2017 On Geometric Features for Skeleton-Based Action Recognition Using Multilayer LSTM Networks
abstract
RNN-based approaches have achieved outstanding performance on action recognition with skeleton inputs. Currently these methods limit their inputs to coordinates of joints and improve the accuracy mainly by extending RNN models to spatial domains in various ways. While such models explore relations between different parts directly from joint coordinates, we provide a simple universal spatial modeling method perpendicular to the RNN model enhancement. Specifically, we select a set of simple geometric features, motivated by the evolution of previous work. With experiments on a 3-layer LSTM framework, we observe that the geometric relational features based on distances between joints and selected lines outperform other features and achieve state-of-art results on four datasets. Further, we show the sparsity of input gate weights in the first LSTM layer trained by geometric features and demonstrate that utilizing joint-line distances as input require less data for training.
Songyang Zhang 0004, Xiaoming Liu 0002, Jun Xiao 0001
WACV3
2017 Disambiguating named entities with deep supervised learning via crowd labels
abstract
Named entity disambiguation (NED) is the task of linking mentions of ambiguous entities to their referenced entities in a knowledge base such as Wikipedia. We propose an approach to effectively disentangle the discriminative features in the manner of collaborative utilization of collective wisdom (via human-labeled crowd labels) and deep learning (via human-generated data) for the NED task. In particular, we devise a crowd model to elicit the underlying features (crowd features) from crowd labels that indicate a matching candidate for each mention, and then use the crowd features to fine-tune a dynamic convolutional neural network (DCNN). The learned DCNN is employed to obtain deep crowd features to enhance traditional hand-crafted features for the NED task. The proposed method substantially benefits from the utilization of crowd knowledge (via crowd labels) into a generic deep learning for the NED task. Experimental analysis demonstrates that the proposed approach is superior to the traditional hand-crafted features when enough crowd labels are gathered.
Le-kui Zhou, Siliang Tang, Jun Xiao 0001, Fei Wu 0001, Yueting Zhuang
Frontiers Inf. Technol. Electron. Eng.3
2017 A human motion feature based on semi-supervised learning of GMM
Qi Tian 0001, Yinfu Feng, Jun Xiao 0001, Hanzhi Zhang, Yueting Zhuang, Xiaosong Yang, Jian J. Zhang 0001
Multim. Syst.3
2017 Hierarchical Contextual Attention Recurrent Neural Network for Map Query Suggestion
abstract
The query logs from an on-line map query system provide rich cues to understand the behaviors of human crowds. With the growing ability of collecting large scale query logs, the query suggestion has been a topic of recent interest. In general, query suggestion aims at recommending a list of relevant queries w.r.t. users’ inputs via an appropriate learning of crowds’ query logs. In this paper, we are particularly interested in map query suggestions (e.g., the predictions of location-related queries) and propose a novel modelHierarchical Contextual Attention Recurrent Neural Network(HCAR-NN) for map query suggestion in an encoding-decoding manner. Given crowds map query logs, our proposed HCAR-NN not only learns the local temporal correlation among map queries in a query session (e.g., queries in a short-term interval are relevant to accomplish a search mission), but also captures the global longer range contextual dependencies among map query sessions in query logs (e.g., how a sequence of queries within a short-term interval has an influence on another sequence of queries). We evaluate our approach over millions of queries from a commercial search engine (i.e.,Baidu Map). Experimental results show that the proposed approach provides significant performance improvements over the competitive existing methods in terms of classical metrics (i.e.,Recall@KandMRR) as well as the prediction of crowds’ search missions.
Jun Song 0004, Jun Xiao 0001, Fei Wu 0001, Haishan Wu, Tong Zhang 0001, Zhongfei Zhang, Wenwu Zhu 0001
IEEE Trans. Knowl. Data Eng.2
2017 Temporal Interaction and Causal Influence in Community-Based Question Answering
abstract
During the last decade, community-based question answering (CQA) sites have accumulated a vast amount of questions and their crowdsourced answers over time. How to efficiently identify the quality of answers that are relevant to a given question has become an active line of research in CQA. The major challenge of CQA is the accurate selection of high-quality answers w.r.t given questions. Previous approaches tend to model the semantic matching between individual pair of one question and its corresponding answer (how fitting an answer is to a posted question). However, these works ignore the temporal interactions between answers (how previous answers influence the late posted answers). For example, a rational user likely adapts others' opinions, revises his inclinations, and posts a more appropriate answer after understanding the given question and previously posted answers. As a result, this paper devises an architecture named Temporal Interaction and Causal Influence LSTM (TC-LSTM) to effectively leverage not only the causal influence between question-answer (how appropriate an answer is for a given question) but also the temporal interactions between answers-answer (how a high-quality answer gradually forms). In particular, long short-term memory (LSTM) is used to capture the explicit question-answer influence and the implicit answers-answer interactions. Experiments are conducted on SemEval 2015 CQA dataset for answer classification task and Baidu Zhidao Dataset for answer ranking task. The experimental results show the advantage of our model comparing with other state-of-the-art methods.
Fei Wu 0001, Xinyu Duan, Jun Xiao 0001, Zhou Zhao 0001, Siliang Tang, Yin Zhang 0006, Yueting Zhuang
IEEE Trans. Knowl. Data Eng.3
2017 Bag-of-Discriminative-Words (BoDW) Representation via Topic Modeling
abstract
Many of the words in a given document either deliver facts (objective) or express opinions (subjective), respectively, depending on the topics they are involved in. For example, given a bunch of documents, the word “bug” assigned to the topic “order Hemiptera” apparently remarks one object (i.e., one kind of insects), while the same word assigned to the topic “software” probably conveys a negative opinion. Motivated by the intuitive assumption that different words have varying degrees of discriminative power in delivering the objective sense or the subjective sense with respect to their assigned topics, a model named as discriminatively objective-subjective LDA (dosLDA) is proposed in this paper. The essential idea underlying the proposed dosLDA is that a pair of objective and subjective selection variables are explicitly employed to encode the interplay between topics and discriminative power for the words in documents in a supervised manner. As a result, each document is appropriately represented as “bag-of-discriminativewords” (BoDW). The experiments reported on documents and images demonstrate that dosLDA not only performs competitively over traditional approaches in terms of topic modeling and document classification, but also has the ability to discern the discriminative power of each word in terms of its objective or subjective sense with respect to its assigned topic.
Yueting Zhuang, Hanqi Wang, Jun Xiao 0001, Fei Wu 0001, Yi Yang 0001, Weiming Lu 0001, Zhongfei Zhang
IEEE Trans. Knowl. Data Eng.3
2017 Matryoshka Peek: Toward Learning Fine-Grained, Robust, Discriminative Features for Product Search
abstract
In sharp contrast to the traditional category/subcategory level image retrieval, product image search aims to find the images containing the exact same product. This is a challenging problem because in addition to being robust under different imaging conditions such as varying viewpoints and illumination changes, the features should also be able to distinguish the specific product among many similar products. Consequently, it is important to utilize a large dataset, containing many product classes, to learn a strongly discriminative representation. Building such a dataset requires laborious manual annotation. Toward learning fine-grained, robust, discriminative features for product image search, we present a novel paradigm that can construct the required dataset without any human annotation. Unlike other fine-grained recognition works that rely on high-quality annotated datasets and are very narrowly focused on a specific object category, our method handles multiple object classes and requires minimum human effort. First, an ImageNet pretrained model is used to generate product clusters. As the original features from ImageNet are not discriminative, the clusters generated by this unsupervised procedure contain much noise. We alleviate noise by explicitly modeling noise distribution and automatically detecting errors during learning. The proposed paradigm is general, requires minimum human efforts, and is applicable to any deep learning task where fine-grained discriminative features are desired. Extensive experiments on the ALISC dataset have demonstrated that our approach is sound and effective, surpassing the baseline GoogleNet model by 15.09%.
Zawlin Kyaw, Shuhan Qi, Ke Gao 0012, Hanwang Zhang, Jun Xiao 0001, Xuan Wang 0002, Tat-Seng Chua
IEEE Trans. Multim.6
2016 A 3D human motion refinement method based on sparse motion bases selection
abstract
Motion capture (MOCAP) is an important technique that is widely used in many areas such as computer animation, film industry, physical training and so on. Even with professional MOCAP system, the missing marker problems always occur. Motion refinement is an essential preprocessing step for MOCAP data based applications. Although many existing approaches for motion refinement have been developed, it is still a challenging task due to the complexity and diversity of human motion. A data driven based motion refinement method is proposed in this paper, which modifies the traditional sparse coding process for special task of motion recovery from missing parts. Meanwhile, the objective function is derived by taking both statistical and kinematical property of motion data into account. Poselet model and moving window grouping are applied in the proposed method to achieve a fine-grained feature representation, which preserves the embedded spatial-temporal kinematic information. 5 motion dictionaries are learnt for each kind of poselet from training data in parallel. The motion refine problem is finally solved as an ℓ1-minimization problem. Compared with several state-of-art motion refine methods, the experimental result shows that our approach outperforms the competitors.
Yinfu Feng, Shuang Liu 0006, Jun Xiao 0001, Xiaosong Yang, Jian J. Zhang 0001
CASA4
2016 Self-Paced Boost Learning for Classification
Te Pi, Xi Li 0001, Zhongfei Zhang, Deyu Meng, Fei Wu 0001, Jun Xiao 0001, Yueting Zhuang
IJCAI6
2016 Diverse Image Captioning via GroupTalk
Zhuhao Wang, Fei Wu 0001, Weiming Lu 0001, Jun Xiao 0001, Xi Li 0001, Yueting Zhuang
IJCAI4
2016 LSTM-in-LSTM for generating long descriptions of images
abstract
In this paper, we propose an approach for generating rich fine-grained textual descriptions of images. In particular, we use an LSTM-in-LSTM (long short-term memory) architecture, which consists of an inner LSTM and an outer LSTM. The inner LSTM effectively encodes the long-range implicit contextual interaction between visual cues (i.e., the spatiallyconcurrent visual objects), while the outer LSTM generally captures the explicit multi-modal relationship between sentences and images (i.e., the correspondence of sentences and images). This architecture is capable of producing a long description by predicting one word at every time step conditioned on the previously generated word, a hidden vector (via the outer LSTM), and a context vector of fine-grained visual cues (via the inner LSTM). Our model outperforms state-of-theart methods on several benchmark datasets (Flickr8k, Flickr30k, MSCOCO) when used to generate long rich fine-grained descriptions of given images in terms of four different metrics (BLEU, CIDEr, ROUGE-L, and METEOR).
Jun Song 0004, Siliang Tang, Jun Xiao 0001, Fei Wu 0001, Zhongfei Zhang
Comput. Vis. Media3
2016 Fast view-based 3D model retrieval via unsupervised multiple feature fusion and online projection learning
Jun Xiao 0001, Yinfu Feng, Mingming Ji, Yueting Zhuang
Signal Process.1
2016 Structure-Aware Slow Feature Analysis for Age Estimation
abstract
As an important and challenging problem in computer vision, face age estimation is typically cast as a classification or regression problem over a set of face samples. However, most existing efforts to age estimation usually cope with the face samples individually, which do not take full advantage of the temporal structure and contextual structure of the face samples. In this letter, we propose an age estimation approach named structure-aware slow feature analysis, which is capable of effectively capturing the structure of human faces in the aspects of time-related smoothness for progressive age variation as well as face-related attribute constraints for face age consistency. As a result, we present an iterative optimization scheme to effectively learn the slowly varying feature transformation. Experimental results demonstrate the effectiveness of our approach on the Morph dataset.
Zhouzhou He, Xi Li 0001, Zhongfei Zhang, Jun Xiao 0001
IEEE Signal Process. Lett.5
2015 Metric Learning Driven Multi-Task Structured Output Optimization for Robust Keypoint Tracking
abstract
As an important and challenging problem in computer vision and graphics, keypoint-based object tracking is typically formulated in a spatio-temporal statistical learning framework. However, most existing keypoint trackers are incapable of effectively modeling and balancing the following three aspects in a simultaneous manner: temporal model coherence across frames, spatial model consistency within frames, and discriminative feature construction. To address this issue, we propose a robust keypoint tracker based on spatio-temporal multi-task structured output optimization driven by discriminative metric learning. Consequently, temporal model coherence is characterized by multi-task structured keypoint model learning over several adjacent frames, while spatial model consistency is modeled by solving a geometric verification based structured learning problem. Discriminative feature construction is enabled by metric learning to ensure the intra-class compactness and inter-class separability. Finally, the above three modules are simultaneously optimized in a joint learning scheme. Experimental results have demonstrated the effectiveness of our tracker.
Xi Li 0001, Jun Xiao 0001, Fei Wu 0001, Yueting Zhuang
AAAI3
2015 Continuous Angle-based Outlier Detection on High-dimensional Data Streams
abstract
Outlier detection over data streams is an increasingly important task in data mining. Traditional distance-based data stream outlier detection is unsuitable for high-dimensional data sets, since the discrimination of distances between different data points becomes rather poor in high dimensional space. ABOD (Angle-based Outlier Detection) is an effective approach to detecting outliers in high-dimensional space. In this paper, the problem of continuous ABOD over data streams is studied. Generally, only a few data objects may change their states during two consecutive timestamps. Therefore, we propose several incremental angle-based outlier detection approaches over data streams based on ABOD and its variants that provide visible speed-up without loss of accuracy. Firstly, the basic ideas of these incremental algorithms are introduced. Then, we explain the time complexity of them. Finally, we use synthetic data streams to prove their efficiency.
Hiroyuki Kitagawa, Jun Xiao 0001
IDEAS3
2015 A locally weighted sparse graph regularized Non-Negative Matrix Factorization method
Yinfu Feng, Jun Xiao 0001, Yueting Zhuang
Neurocomputing2
2015 Efficient semi-supervised multiple feature fusion with out-of-sample extension for 3D model retrieval
Mingming Ji, Yinfu Feng, Jun Xiao 0001, Yueting Zhuang, Xiaosong Yang, Jian J. Zhang 0001
Neurocomputing3
2015 View-invariant human action recognition via robust locally adaptive multi-view learning
abstract
Human action recognition is currently one of the most active research areas in computer vision. It has been widely used in many applications, such as intelligent surveillance, perceptual interface, and content-based video retrieval. However, some extrinsic factors are barriers for the development of action recognition; e.g., human actions may be observed from arbitrary camera viewpoints in realistic scene. Thus, view-invariant analysis becomes important for action recognition algorithms, and a number of researchers have paid much attention to this issue. In this paper, we present a multi-view learning approach to recognize human actions from different views. As most existing multi-view learning algorithms often suffer from the problem of lacking data adaptiveness in the nearest neighborhood graph construction procedure, a robust locally adaptive multi-view learning algorithm based on learning multiple local L1-graphs is proposed. Moreover, an efficient iterative optimization method is proposed to solve the proposed objective function. Experiments on three public view-invariant action recognition datasets, i.e., ViHASi, IXMAS, and WVU, demonstrate data adaptiveness, effectiveness, and efficiency of our algorithm. More importantly, when the feature dimension is correctly selected (i.e., >60), the proposed algorithm stably outperforms state-of-the-art counterparts and obtains about 6% improvement in recognition accuracy on the three datasets.
Jia-geng Feng, Jun Xiao 0001
Frontiers Inf. Technol. Electron. Eng.2
2015 Sparse motion bases selection for human motion denoising
abstract
Human motion denoising is an indispensable step of data preprocessing for many motion data based applications. In this paper, we propose a data-driven based human motion denoising method that sparsely selects the most correlated subset of motion bases for clean motion reconstruction. Meanwhile, it takes the statistic property of two common noises, i.e., Gaussian noise and outliers, into account in deriving the objective functions. In particular, our method firstly divides each human pose into five partitions termed as poselets to gain a much fine-grained pose representation. Then, these poselets are reorganized into multiple overlapped poselet groups using a lagged window moving across the entire motion sequence to preserve the embedded spatial–temporal motion patterns. Afterward, five compacted and representative motion dictionaries are constructed in parallel by means of fast K-SVD in the training phase; they are used to remove the noise and outliers from noisy motion sequences in the testing phase by solving ℓ 1 -minimization problems. Extensive experiments show that our method outperforms its competitors. More importantly, compared with other data-driven based method, our method does not need to specifically choose the training data , it can be more easily applied to real-world applications.
Jun Xiao 0001, Yinfu Feng, Mingming Ji, Xiaosong Yang, Jian J. Zhang 0001, Yueting Zhuang
Signal Process.1
2015 Sketch-based human motion retrieval via selected 2D geometric posture descriptor
abstract
Sketch-based human motion retrieval is a hot topic in computer animation in recent years. In this paper, we present a novel sketch-based human motion retrieval method via selected 2-dimensional (2D) Geometric Posture Descriptor (2GPD). Specially, we firstly propose a rich 2D pose feature call 2D Geometric Posture Descriptor (2GPD), which is effective in encoding the 2D posture similarity by exploiting the geometric relationships among different human body parts. Since the original 2GPD is of high dimension and redundant, a semi-supervised feature selection algorithm derived from Laplacian Score is then adopted to select the most discriminative feature component of 2GPD as feature representation, and we call it as selected 2GPD. Finally, a posture-by-posture motion retrieval algorithm is used to retrieve a motion sequence by sketching several key postures. Experimental results on CMU human motion database demonstrate the effectiveness of our proposed approach.
Jun Xiao 0001, Zhangpeng Tang, Yinfu Feng, Zhidong Xiao
Signal Process.1
2015 Mining Spatial-Temporal Patterns and Structural Sparsity for Human Motion Data Denoising
abstract
Motion capture is an important technique with a wide range of applications in areas such as computer vision, computer animation, film production, and medical rehabilitation. Even with the professional motion capture systems, the acquired raw data mostly contain inevitable noises and outliers. To denoise the data, numerous methods have been developed, while this problem still remains a challenge due to the high complexity of human motion and the diversity of real-life situations. In this paper, we propose a data-driven-based robust human motion denoising approach by mining the spatial-temporal patterns and the structural sparsity embedded in motion data. We first replace the regularly used entire pose model with a much fine-grained partlet model as feature representation to exploit the abundant local body part posture and movement similarities. Then, a robust dictionary learning algorithm is proposed to learn multiple compact and representative motion dictionaries from the training data in parallel. Finally, we reformulate the human motion denoising problem as a robust structured sparse coding problem in which both the noise distribution information and the temporal smoothness property of human motion have been jointly taken into account. Compared with several state-of-the-art motion denoising methods on both the synthetic and real noisy motion data, our method consistently yields better performance than its counterparts. The outputs of our approach are much more stable than that of the others. In addition, it is much easier to setup the training dataset of our method than that of the other data-driven-based methods.
Yinfu Feng, Mingming Ji, Jun Xiao 0001, Xiaosong Yang, Jian J. Zhang 0001, Yueting Zhuang, Xuelong Li 0001
IEEE Trans. Cybern.3
2014 Exploiting temporal stability and low-rank structure for motion capture data refinement
abstract
Inspired by the development of the matrix completion theories and algorithms, a low-rank based motion capture (mocap) data refinement method has been developed, which has achieved encouraging results. However, it does not guarantee a stable outcome if we only consider the low-rank property of the motion data. To solve this problem, we propose to exploit the temporal stability of human motion and convert the mocap data refinement problem into a robust matrix completion problem, where both the low-rank structure and temporal stability properties of the mocap data as well as the noise effect are considered. An efficient optimization method derived from the augmented Lagrange multiplier algorithm is presented to solve the proposed model. Besides, a trust data detection method is also introduced to improve the degree of automation for processing the entire set of the data and boost the performance. Extensive experiments and comparisons with other methods demonstrate the effectiveness of our approaches on both predicting missing data and de-noising.
Yinfu Feng, Jun Xiao 0001, Yueting Zhuang, Xiaosong Yang, Jian J. Zhang 0001, Rong Song
Inf. Sci.2
2014 Human motion retrieval based on freehand sketch
abstract
ABSTRACT In this paper, we present an integrated framework of human motion retrieval based on freehand sketch. With some simple rules, the user can acquire a desired motion by sketching several key postures. To retrieve efficiently and accurately by sketch, the 3D postures are projected onto several 2D planes. The limb direction feature is proposed to represent the input sketch and the projected‐postures. Furthermore, a novel index structure based on k‐d tree is constructed to index the motions in the database, which speeds up the retrieval process. With our posture‐by‐posture retrieval algorithm, a continuous motion can be got directly or generated by using a pre‐computed graph structure. What's more, our system provides an intuitive user interface. The experimental results demonstrate the effectiveness of our method. © 2014 The Authors.Computer Animation and Virtual Worldspublished by John Wiley & Sons, Ltd.
Zhangpeng Tang, Jun Xiao 0001, Yinfu Feng, Xiaosong Yang
Comput. Animat. Virtual Worlds2
2014 Real-time motion data annotation via action string
abstract
ABSTRACT Even though there is an explosive growth of motion capture data, there is still a lack of efficient and reliable methods to automatically annotate all the motions in a database. Moreover, because of the popularity of mocap devices in home entertainment systems, real‐time human motion annotation or recognition becomes more and more imperative. This paper presents a new motion annotation method that achieves both the aforementioned two targets at the same time. It uses a probabilistic pose feature based on the Gaussian Mixture Model to represent each pose. After training a clustered pose feature model, a motion clip could be represented as an action string. Then, a dynamic programming‐based string matching method is introduced to compare the differences between action strings. Finally, in order to achieve the real‐time target, we construct a hierarchical action string structure to quickly label each given action string. The experimental results demonstrate the efficacy and efficiency of our method. Copyright © 2014 John Wiley & Sons, Ltd.
Qi Tian 0001, Jun Xiao 0001, Yueting Zhuang, Hanzhi Zhang, Xiaosong Yang, Jian J. Zhang 0001, Yinfu Feng
Comput. Animat. Virtual Worlds2
2013 Hypergraph Spectral Hashing for image retrieval with heterogeneous social contexts
Yang Liu 0098, Jian Shao 0001, Jun Xiao 0001, Fei Wu 0001, Yueting Zhuang
Neurocomputing3
2013 A semantic feature for human motion retrieval
abstract
ABSTRACT With the explosive growth of motion capture data, it becomes very imperative in animation production to have an efficient search engine to retrieve motions from large motion repository. However, because of the high dimension of data space and complexity of matching methods, most of the existing approaches cannot return the result in real time. This paper proposes a high level semantic feature in a low dimensional space to represent the essential characteristic of different motion classes. On the basis of the statistic training of Gauss Mixture Model, this feature can effectively achieve motion matching on both global clip level and local frame level. Experiment results show that our approach can retrieve similar motions with rankings from large motion database in real‐time and also can make motion annotation automatically on the fly. Copyright © 2013 John Wiley & Sons, Ltd.
Qi Tian 0001, Yinfu Feng, Jun Xiao 0001, Yueting Zhuang, Xiaosong Yang, Jian J. Zhang 0001
Comput. Animat. Virtual Worlds3
2013 Retrieval-based cartoon gesture recognition and applications via semi-supervised heterogeneous classifiers learning
Zhang Liang, Yueting Zhuang, Yi Yang 0001, Jun Xiao 0001
Pattern Recognit.4
2012 Adaptive Unsupervised Multi-view Feature Selection for Visual Concept Recognition
Yinfu Feng, Jun Xiao 0001, Yueting Zhuang, Xiaoming Liu 0002
ACCV (1)2
2012 Active learning for social image retrieval using Locally Regressive Optimal Design
Yinfu Feng, Jun Xiao 0001, Zhengjun Zha, Yi Yang 0001
Neurocomputing2
2012 Synthesizing style-preserving cartoons via non-negative style factorization
abstract
We present a complete framework for synthesizing style-preserving 2D cartoons by learning from traditional Chinese cartoons. In contrast to reusing-based approaches which rely on rearranging or retrieving existing cartoon sequences, we aim to generate stylized cartoons with the idea of style factorization. Specifically, starting with 2D skeleton features of cartoon characters extracted by an improved rotoscoping system, we present a non-negative style factorization (NNSF) algorithm to obtain style basis and weights and simultaneously preserve class separability. Thus, factorized style basis can be combined with heterogeneous weights to re-synthesize style-preserving features, and then these features are used as the driving source in the character reshaping process via our proposed subkey-driving strategy. Extensive experiments and examples demonstrate the effectiveness of the proposed framework.
Zhang Liang, Jun Xiao 0001, Yueting Zhuang
J. Zhejiang Univ. Sci. C2
2011 Predicting missing markers in human motion capture using l1-sparse representation
abstract
Abstract Missing marker problem is very common in human motion capture. In contrast to most current methods which handle this problem based on trying to learn a reliable predictor from the observations, we consider it from the perspective of sparse representation and propose a novel method which is namedl1‐sparse representation of missing markers prediction (L1‐SRMMP). We assume that the incomplete pose can be represented by a linear combination of a few poses from the training set and the representation is sparse. Therefore, we cast the predicting missing markers as finding a sparse representation of the observable data of the incomplete pose, and then we use it to predict the missing data. In order to get a sparse representation, we employl1‐norm in our objective function. Moreover, we propose presentation coefficient weighted update (PCWU) algorithm to mitigate the limited capacity problem of the training set. Experimental results demonstrate the effectiveness and efficiency of our method to predict the missing markers in human motion capture. Copyright © 2011 John Wiley & Sons, Ltd.
Jun Xiao 0001, Yinfu Feng, Wenyuan Hu
Comput. Animat. Virtual Worlds1
2011 Learning a 3D Human Pose Distance Metric from Geometric Pose Descriptor
abstract
Estimating 3D pose similarity is a fundamental problem on 3D motion data. Most previous work calculates L2-like distance of joint orientations or coordinates, which does not sufficiently reflect the pose similarity of human perception. In this paper, we present a new pose distance metric. First, we propose a new rich pose feature set called Geometric Pose Descriptor (GPD). GPD is more effective in encoding pose similarity by utilizing features on geometric relations among body parts, as well as temporal information such as velocities and accelerations. Based on GPD, we propose a semisupervised distance metric learning algorithm called Regularized Distance Metric Learning with Sparse Representation (RDSR), which integrates information from both unsupervised data relationship and labels. We apply the proposed pose distance metric to applications of motion transition decision and content-based pose retrieval. Quantitative evaluations demonstrate that our method achieves better results with only a small amount of human labels, showing that the proposed pose distance metric is a promising building block for various 3D-motion related applications.
Cheng Chen 0023, Yueting Zhuang, Feiping Nie 0001, Yi Yang 0001, Fei Wu 0001, Jun Xiao 0001
IEEE Trans. Vis. Comput. Graph.6
2010 Silhouette representation and matching for 3D pose discrimination - A comparative study
Cheng Chen 0023, Yueting Zhuang, Jun Xiao 0001
Image Vis. Comput.3
2010 A group of novel approaches and a toolkit for motion capture data reusing
Jun Xiao 0001, Yueting Zhuang, Fei Wu 0001, Tongqiang Guo, Zhang Liang
Multim. Tools Appl.1
2009 Perceptual 3D pose distance estimation by boosting relational geometric features
abstract
Abstract Traditional pose similarity functions based on joint coordinates or rotations often do not conform to human perception. We propose a new perceptual pose distance:Relational Geometric Distancethat accumulates the differences over a set of features that reflects the geometric relations between different body parts. An extensive relational geometric feature pool that contains a large number of potential features is defined, and the features effective for pose similarity estimation are selected using a set of labeled data by Adaboost. The extensive feature pool guarantees that a wide diversity of features is considered, and the boosting ensures that the selected features are optimized when used jointly. Finally, the selected features form a pose distance function that can be used for novel poses. Experiments show that our method outperforms others in emulating human perception in pose similarity. Our method can also adapt to specific motion types and capture the features that are important for pose similarity of a certain motion type. Copyright © 2009 John Wiley & Sons, Ltd.
Cheng Chen 0023, Yueting Zhuang, Jun Xiao 0001, Zhang Liang
Comput. Animat. Virtual Worlds3
2009 Competitive motion synthesis based on hybrid control
abstract
Abstract We propose a simple and effective framework to deal with the problem of synthesizing interactive and competitive motions while reflecting the interactions. Two nontrivial issues are addressed in synthesizing two‐character motions in competitive environment: how to reveal the embedded routines while keeping visual reality and how to build interactive models based on singly captured motions. To solve these issues, we employ a hierarchical framework: the finite state machine (FSM) controls the state transition in the higher layer, and the hybrid approach controls the action selection in the lower layer. The proposed approach contains two folds: first, a rule‐based control scheme is proposed to simulate routine steps based on statistical analysis. Second, the interactive models are designed for simulating dense interactions between two players. The Relevance Vector Machine (RVM) algorithm is adopted to select attack styles and coupled with motion transition graph to determine combination blows. Here we apply the proposed framework of hybrid paradigm to boxing sport as an example. Copyright © 2009 John Wiley & Sons, Ltd.
Liang Zhang 0045, Jun Xiao 0001, Yueting Zhuang, Cheng Chen 0023
Comput. Animat. Virtual Worlds2
2008 Adaptive and compact shape descriptor by progressive feature combination and selection with boosting
abstract
Many types of shape descriptors have been proposed for 2D shape analysis, but most of them consist of component features that are not adapted to specific problems. This has two drawbacks. First, computation is wasted on the irrelevant components; second, the accuracy is impaired. This paper proposes an effective method that generates compact descriptors adapted to specific problems in hand, where each component of the new descriptor is a linear combination of the components in some classic descriptors. A progressive strategy is used to construct and select the most suitable linear combinations in successive rounds, where a variant of Adaboost is employed to ensure the optimum of the selected combinations in each round. Experiments show that our method effectively generates adaptive and compact descriptors for typical applications such as shape classification and retrieval.
Cheng Chen 0023, Yueting Zhuang, Jun Xiao 0001, Fei Wu 0001
CVPR3
2008 Active post-refined multimodality video semantic concept detection with tensor representation
abstract
In this paper, we resolve the problem of multi-modality video representation and semantic concept detection. Interaction and integration of multi-modality media types such as visual, audio and textual data in video are essential to video semantic analysis. Traditionally, videos are represented as vectors in the Euclidean space. Many learning algorithms are then taken to these vectors in a high dimensional space for dimension reduction, classification, clustering and so on. However, the multiple modalities in video not only have their own properties, but also have correlations among them; whereas the simple vector representation weakens the power of these relatively independent modalities and even ignores their relations to some extent. In this paper, we introduce a higher-order tensor framework for video analysis, in which we represent image, video and text three modalities in video shots as data points by the 3rd-order tensor called tensorshots. We propose a novel dimension reduction method that explicitly considers the manifold structure of the tensor space from multimodal media data which is temporal associated co-occurrence and then detect video semantic concepts through powerful classifiers which take tensor as input. Our algorithm preserves the intrinsic structure of the submanifold where tensorshots are sampled, and is also able to map out-of-sample data points directly. Moreover we apply an active learning based contextual and temporal post-refining strategy to enhance detection accuracy. Experiment results show that our method improves the performance of video semantic concept detection.
Fei Wu 0001, Yueting Zhuang, Jun Xiao 0001
ACM Multimedia4
2008 Perspective-aware cartoon clips synthesis
abstract
Abstract In this paper we propose an approach, which allows the users to synthesize cartoon clips according to the perspective of the background image. In order to construct the cartoons smoothly, the character's edge distance and motion direction distance are demonstrated to be the factors affecting the human perception in similarity evaluation, and utilized in cartoon clips synthesis. When applying the generated cartoons to the background image, in which the perspective exists, the size of the character is coordinated according to the scaling factor calculated from the vanishing line. The experiment results demonstrate that our approach can synthesize the cartoon clips more smoothly compared with other single frame reusing strategies. The generated cartoons, which are applied to the background image, can be accepted by the human perception well. Copyright © 2008 John Wiley & Sons, Ltd.
Yueting Zhuang, Jun Yu 0002, Jun Xiao 0001, Cheng Chen 0023
Comput. Animat. Virtual Worlds3
2007 Adaptive control in cartoon data reusing
abstract
Abstract In this paper, we propose a novel approach, which reuses Traditional Chinese Cartoon to create new animations. In order to extract the cartoon character precisely, a segmentation method based on edge detection is implemented. Before reusing the data, a lower‐ dimensional space of the cartoon data is constructed by ISOmap. The character's gesture difference calculated by optical flow is combined with character's edge difference through a novel distance function, which is controlled by a weight parameter. The animation is created by reordering the existing data into a sequence, which is the shortest path between two designated data in the space. Our approach utilizes image processing, computer vision, and machine learning in cartoon creation and the experiment results demonstrate that the animation's quality can be effectively improved by the fusion of these techniques. Copyright © 2007 John Wiley & Sons, Ltd.
Jun Yu 0002, Yueting Zhuang, Jun Xiao 0001, Cheng Chen 0023
Comput. Animat. Virtual Worlds3
2006 An Efficient Keyframe Extraction from Motion Capture Data
Jun Xiao 0001, Yueting Zhuang, Fei Wu 0001
Computer Graphics International1
2005 Automatic generation of human animation based on motion programming
abstract
Abstract In motion simulations, video games and animation films, lots of interactions between characters and virtual environments are needed. Even though realistic motion data can be derived from MoCap system, motion editing and synthesis, animators must adapt these motion data to specific virtual environment manually, which is a boring and time‐consuming job. Here we propose a framework to program the movements of characters and generate navigation animations in virtual environment. Given a virtual environment, a visual user interface is provided for animators to interactively generate motion scripts, describing the characters' movements in this scene and finally used to retrieve motion clips from MoCap database and generate navigation animations automatically. This framework also provides flexible mechanism for animators to get varied resulting animations by configurable table of motion bias coefficients and interactive visual user interface. Copyright © 2005 John Wiley & Sons, Ltd.
Yueting Zhuang, Jun Xiao 0001, Yizi Wu, Fei Wu 0001
Comput. Animat. Virtual Worlds2