EDBT 2026 Demo / reviewers in the wild / expert
Muli Yang
dblp:233/9785
· DBLP profile ↗
35ranked-venue papers
9as first author
29since 2021 · last 2026
0000-0002-5959-2931ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 8 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 6 first-author · 18 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Illumination-Aware Restoration of Metalens-Captured Images: A New Dataset and a Strong BaselineabstractMetalenses offer compelling advantages such as lightweight and ultra-thin design, making them promising alternatives to conventional lenses. However, their widespread adoption is hindered by image quality degradation caused by chromatic and angular aberrations. To mitigate this, restoration processes are often necessary to recover high-quality RGB images from metalens-captured inputs. While recent deep learning-based restoration methods show promise, they typically (1) blur or distort peripheral regions, or (2) fail entirely under unseen illumination conditions. To advance metalens image restoration, we introduce IlluMeta---the first and largest real-world, illumination-aware metalens image dataset—captured across diverse lighting environments. In addition, we propose a novel end-to-end restoration framework that directs attention to challenging regions and adaptively adjusts to varying illuminations via reinforcement learning. Experiments show that our method can be applied in a plug-and-play manner to enhance existing models, significantly improving image restoration quality, especially under unseen lighting conditions, paving the way for broader real-world deployment of metalens technologies. Fen Fang, Xinan Liang, Muli Yang, Jinghong Zheng 0001, Tobias Wilhelm W. Mass, Ying Sun 0001, Xulei Yang, Xuewu Xu, Zhengguo Li |
AAAI | 3 |
| 2026 | Next-Generation Metalens Vision System: Powered by AI and Applied to AIabstractMetalenses have been widely recognized as a key building block of next-generation optical systems, offering unprecedented advantages in compactness, lightweight design, and scalable manufacturing compared to traditional refractive optics. Despite this promise, practical use is limited by optical aberrations, blur, and illumination sensitivity, which degrade both visual quality and machine perception. In this demonstration, we present an end-to-end metalens vision system—from hardware sensing with a custom-built RGB metalens camera, to physics-informed imaging and real-time restoration, and finally to downstream vision applications such as object detection and depth estimation. By integrating spatially-aware attention enhancement and reinforcement learning-based illumination control into a real-time system, our solution transforms degraded raw captures into high-fidelity images that are both visually interpretable and functionally reliable for machine vision. This AI-powered pipeline highlights metalenses as a cornerstone for next-generation imaging, where advances in optics and machine intelligence jointly drive the future of visual perception. Fen Fang, Muli Yang, Henan Wang, Xinan Liang, Tobias Wilhelm W. Mass, Xuewu Xu, Xulei Yang, Zhengguo Li |
AAAI | 2 |
| 2026 | Editing Is a Bargaining Game: Balanced Knowledge Editing in Large Language ModelsabstractLarge Language Models (LLMs) are prone to generating incorrect or outdated information, thereby necessitating efficient and precise mechanisms for knowledge updates. Existing knowledge editing approaches, however, often encounter conflicts between two competing objectives: maintaining existing knowledge (preservation) and incorporating new information (editing). During gradient-based optimization, these conflicting objectives can lead to imbalanced update directions, where one gradient dominates, ultimately resulting in suboptimal learning dynamics. To address this challenge, we propose a balanced knowledge editing framework inspired by Nash bargaining theory. Our method guides the optimization process toward a Pareto stationary point, ensuring an equilibrium solution wherein any deviation from the final state would degrade the overall performance with respect to both objectives. This guarantees optimality in preserving prior knowledge while integrating new information. We empirically validate the effectiveness of our approach across a range of evaluation metrics on standard benchmark datasets. Extensive experiments show that our method consistently outperforms state-of-the-art techniques, achieving a superior balance between knowledge preservation and update accuracy. Jiexi Yan, Muli Yang, Fen Fang, Cheng Deng 0002 |
AAAI | 3 |
| 2026 | Your AI-Generated Image Detector Can Secretly Achieve SOTA Accuracy, If CalibratedabstractDespite being trained on balanced datasets, existing AI-generated image detectors often exhibit systematic bias at test time, frequently misclassifying fake images as real. We hypothesize that this behavior stems from distributional shift in fake samples and implicit priors learned during training. Specifically, models tend to overfit to superficial artifacts that do not generalize well across different generation methods, leading to a misaligned decision threshold when faced with test-time distribution shift. To address this, we propose a theoretically grounded post-hoc calibration framework based on Bayesian decision theory. In particular, we introduce a learnable scalar correction to the model’s logits, optimized on a small validation set from the target distribution while keeping the backbone frozen. This parametric adjustment compensates for distributional shift in model output, realigning the decision boundary even without requiring ground-truth labels. Experiments on challenging benchmarks show that our approach significantly improves robustness without retraining, offering a lightweight and principled solution for reliable and adaptive AI-generated image detection in the open world. Muli Yang, Gabriel James Goenawan, Henan Wang, Huaiyuan Qin, Yanhua Yang, Fen Fang, Ying Sun 0001, Joo-Hwee Lim, Hongyuan Zhu 0002 |
AAAI | 1 |
| 2026 | From Language to Driving: A Dual-Loop SLM-Enhanced Framework for Multi-Planner Scheduling via a Domain-Specific LanguageabstractJiawei Liu, Xun Gong, Muli Yang, Xingrui Yu, Fen Fang, Xulei Yang, Ivor Tsang, Yunfeng hu, Hong Chen, Qing Guo. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xun Gong 0007, Muli Yang, Xingrui Yu, Fen Fang, Xulei Yang, Ivor W. Tsang, Yunfeng Hu 0003, Hong Chen 0003, Qing Guo 0005 |
ACL (1) | 3 |
| 2026 | Fisher-Driven Adaptive Locating for Knowledge Editing in Large Language ModelsabstractLarge language models (LLMs) store extensive factual knowledge acquired during pretraining, yet this knowledge is inherently static and may become inaccurate or outdated, leading to knowledge hallucinations.Knowledge editing offers an efficient alternative to full retraining by enabling targeted factual updates while preserving overall model behavior.Existing locate-then-edit methods, however, rely on fixed layer selection strategies, treating the locating stage as a static design choice and failing to account for the hierarchical and instancedependent nature of knowledge representation in LLMs.In this paper, we propose FiDAL, a Fisher-driven adaptation-aware locating strategy that dynamically identifies which model components should be edited for a given knowledge update.FiDAL formulates localization as a weight-level decision problem and leverages Fisher Information to select layers that are both influential and sensitive to factual modifications.A lightweight probing stage with low-rank modulation enables efficient localization with minimal overhead.Experiments on standard benchmarks demonstrate that FiDAL consistently improves editing effectiveness and knowledge preservation across multiple editing methods. Jiexi Yan, Guangtao Lyu, Muli Yang, Cheng Deng 0002 |
ACL (1) | 5 |
| 2026 | Toward Accurate Procedure Planning in Instructional Videos: Visual State Generation Helps Task-Selective DiffusionabstractProcedure planning in instructional videos entails predicting an action sequence that transitions a given start state to a desired goal state. This task is particularly challenging due to two key sources of uncertainty: limited visual observations and an enormous decision space. The former results in multiple plausible plan variations due to missing intermediate visual states, while the latter complicates prediction by requiring selection from a large set of potential actions. Unlike prior work that addresses these issues implicitly, we propose an explicit solution. To mitigate the first challenge, we employ image generation models to synthesize diverse intermediate visual states using various text prompts, followed by a prompt selection module integrated within a diffusion model. To tackle the second challenge, we introduce a task-selective diffusion model that applies a task-specific mask to constrain the action space. As the effectiveness of this mask depends on accurate task classification, we further enhance visual representation by leveraging pre-trained vision-language models to generate action-aware, text-enriched multimodal embeddings. Extensive experiments on three benchmark datasets validate the superior performance of our proposed approach. Fen Fang, Muli Yang, Min Wu 0008, Yanhua Yang, Qianli Xu, Joo-Hwee Lim, Xulei Yang, Hongyuan Zhu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | Counterfactual Risk Minimization for Out-of-Distribution GeneralizationabstractThe out-of-distribution (OOD) property in data is deemed as one main challenge hindering the generalization ability of machine learning algorithms. However, the underlying reasons for this property remain an intriguing and open question that has yet to be fully understood. In this paper, we seek to enhance our understanding of the OOD phenomenon by framing it as a problem of distribution shift and addressing it through two complementary causal perspectives. The first is a generative causal view that elucidates the data generation process. We introduce a novel three-dimensional coordinate system to represent three fundamental distribution shifts, illustrating their role in various OOD generalization problems. The second is an anti-causal view that focuses on the model learning process. We develop an effective approach dubbed Counterfactual Risk Minimization (CRM) to address arbitrary distribution shifts in a unified framework. Additionally, we introduce a new multi-domain visual recognition dataset called CONA to facilitate further exploration of OOD generalization. We conduct evaluations of CRM alongside several state-of-the-art competitors on four benchmark datasets under the three distribution shifts. The results not only affirm CRM's superiority but also shed light on potential future directions. Code and data: https://github.com/muliyangm/CRM. Yanhua Yang, Muli Yang, Henan Wang, Cheng Deng 0002, Hongyuan Zhu 0002 |
IEEE Trans. Image Process. | 2 |
| 2025 | Reasoning Mamba: Hypergraph-Guided Region Relation Calculating for Weakly Supervised Affordance GroundingabstractThis paper pays attention to Weakly Supervised Affordance Grounding (WSAG) task that aims to train model to identify affordance regions using human-object interaction images and egocentric images without the need for costly pixel-level annotations. Most existing methods usually consider the affordance regions to be isolated and directly employ class activation maps to conduct localization, ignoring the relationships with other object components and weakening the performance. For example, a cup’s handle is combined with its body to achieve the pouring ability. Obviously, capturing the region relationships is beneficial for improving the localization accuracy of affordance regions. To this end, we first explore exploiting hypergraph to discover these relations and propose a Reasoning Mamba (R-Mamba) framework. We first extract feature embeddings from exocentric and egocentric images to construct the hypergraphs consisting of multiple vertices and hyperedges, which capture the in-context local region relationships between different visual components. Subsequently, we design a Hypergraph-Guided State Space (HSS) block to reorganize these local relationships from the global perspective. By this mechanism, the model could leverage the captured relationships to improve the localization accuracy of affordance regions. Extensive experiments and visualization analyses demonstrate the superiority of our method. Aming Wu, Muli Yang, Yukuan Min, Yihang Zhu, Cheng Deng 0002 |
CVPR | 3 |
| 2025 | Detecting Open World Objects via Partial Attribute AssignmentabstractDespite being trained on massive data, today’s vision foundation models still fall short in detecting open world objects. Apart from recognizing known objects from training, a successful Open World Object Detection (OWOD) system must also be able to detect unknown objects never seen before, without confusing them with the backgrounds. Unlike prevailing prior works that rely on probability models to learn "objectness", we focus on learning fine-grained, class-agnostic attributes, allowing the detection of both known and unknown objects in an explainable manner. In this paper, we propose Partial Attribute Assignment (PASS), aiming to automatically select and optimize a small, relevant subset of attributes from a large attribute pool. Specifically, we model attribute selection as a Partial Optimal Transport (POT) problem between known visual objects and the attribute pool, in which more relevant attributes signify more transported mass. PASS follows a curriculum schedule that progressively selects and optimizes a targeted subset of attributes during training, promoting stability and accuracy. Our method enjoys end-to-end optimization by minimizing the POT distance and the classification loss on known visual objects, demonstrating high training efficiency and superior OWOD performance among extensive experimental evaluations.‡ Muli Yang, Gabriel James Goenawan, Huaiyuan Qin, Xi Peng 0001, Yanhua Yang, Hongyuan Zhu 0002 |
CVPR | 1 |
| 2025 | Vision-Language Interactive Relation Mining for Open-Vocabulary Scene Graph Generation
Yukuan Min, Muli Yang, Aming Wu, Cheng Deng 0002 |
ICCV | 2 |
| 2025 | Towards Unified Human Motion-Language Understanding via Sparse Interpretable CharacterizationabstractRecently, the comprehensive understanding of human motion has been a prominent area of research due to its critical importance in many fields. However, existing methods often prioritize specific downstream tasks and roughly align text and motion features within a CLIP-like framework. This results in a lack of rich semantic information which restricts a more profound comprehension of human motions, ultimately leading to unsatisfactory performance.
Therefore, we propose a novel motion-language representation paradigm to enhance the interpretability of motion representations by constructing a universal motion-language space, where both motion and text features are concretely lexicalized, ensuring that each element of features carries specific semantic meaning.
Specifically, we introduce a multi-phase strategy mainly comprising Lexical Bottlenecked Masked Language Modeling to enhance the language model's focus on high-entropy words crucial for motion semantics, Contrastive Masked Motion Modeling to strengthen motion feature extraction by capturing spatiotemporal dynamics directly from skeletal motion, Lexical Bottlenecked Masked Motion Modeling to enable the motion model to capture the underlying semantic features of motion for improved cross-modal understanding, and Lexical Contrastive Motion-Language Pretraining to align motion and text lexicon representations, thereby ensuring enhanced cross-modal coherence.
Comprehensive analyses and extensive experiments across multiple public datasets demonstrate that our model achieves state-of-the-art performance across various tasks and scenarios. Guangtao Lyu, Jiexi Yan, Muli Yang, Cheng Deng 0002 |
ICLR | 4 |
| 2025 | Smooth and Flexible Camera Movement Synthesis via Temporal Masked Generative ModelingabstractIn dance performances, choreographers define the visual expression of movement, while cinematographers shape its final presentation through camera work. Consequently, the synthesis of camera movements informed by both music and dance has garnered increasing research interest. While recent advancements have led to notable progress in this area, existing methods predominantly operate in an offline manner—that is, they require access to the entire dance sequence before generating corresponding camera motions. This constraint renders them impractical for real-time applications, particularly in live stage performances, where immediate responsiveness is essential. To address this limitation, we introduce a more practical yet challenging task: online camera movement synthesis, in which camera trajectories must be generated using only the current and preceding segments of dance and music. In this paper, we propose TemMEGA (Temporal Masked Generative Modeling), a unified framework capable of handling both online and offline camera movement generation. TemMEGA consists of three key components. First, a discrete camera tokenizer encodes camera motions as discrete tokens via a discrete quantization scheme. Second, a consecutive memory encoder captures historical context by jointly modeling long- and short-term temporal dependencies across dance and music sequences. Finally, a temporal conditional masked transformer is employed to predict future camera motions by leveraging masked token prediction. Extensive experimental evaluations demonstrate the effectiveness of our TemMEGA, highlighting its superiority in both online and offline camera movement synthesis. Guangtao Lyu, Jiexi Yan, Muli Yang, Cheng Deng 0002 |
NeurIPS | 4 |
| 2025 | Dynamic Adapter Tuning for Long-Tailed Class-Incremental LearningabstractLong-tailed class-incremental learning (LT-CIL) aims to learn new classes continuously from a long-tailed data stream, while simultaneously dealing with challenges such as imbalanced learning of tail classes and catastrophic for-getting. To address these challenges, most existing methods employ a two-stage strategy by initializing model training from scratch with further balanced knowledge driven cali-bration. This strategy faces challenges in deriving discrim-inative features from cold-started backbones for the long-tailed distribution of data, consequently leading to relatively diminished performance. In this paper, with the pow-erful feature extraction capability of pre-trained foundation models, we have achieved a one-stage approach that de-livers superior performance. Specifically, we propose Dy-namic Adapter Tuning (DAT), which employs a dynamic adapter cache mechanism to adapt a pre-trained model to learn tasks sequentially. The adapter in the cache is either dynamically selected or created according to task similar-ity, and further compactified with the new task's adapter to mitigate cross-task and cross-class gaps in LT-CIL, sig-nificantly alleviating catastrophic forgetting and imbalance learning issues, respectively. With extensive experimental validation, our method consistently achieves state-of-the-art performance under the challenging LT-CIL setting. Yanan Gu, Muli Yang, Xu Yang 0019, Hongyuan Zhu 0002, Gabriel James Goenawan, Cheng Deng 0002 |
WACV | 2 |
| 2025 | Consistent Prompt Tuning for Generalized Category Discovery
Muli Yang, Yanan Gu, Cheng Deng 0002, Hanwang Zhang, Hongyuan Zhu 0002 |
Int. J. Comput. Vis. | 1 |
| 2025 | Correction: Consistent Prompt Tuning for Generalized Category Discovery
Muli Yang, Yanan Gu, Cheng Deng 0002, Hanwang Zhang, Hongyuan Zhu 0002 |
Int. J. Comput. Vis. | 1 |
| 2025 | CLIP-based autonomous visual prompting for unsupervised domain incremental learning
Jiaping Yu, Muli Yang, Jiexi Yan, Cheng Deng 0002 |
Neurocomputing | 2 |
| 2025 | Progressive Invariant Causal Feature Learning for Single Domain GeneralizationabstractSingle domain generalization (SDG) aims to transfer models trained on a single source domain to multiple unseen target domains while against the unknown domain shifts. The main challenge lies in learning the domain-invariant features to mitigate the domain shift impact. To address this challenge, we reconsider SDG from a causal perspective to capture the domain-invariant features accurately. Specifically, we present a Progressive Invariant Causal Feature Learning (PICF) method that leverages front-door adjustment to gradually obtain the invariant causal features for SDG. First, we introduce a foreground feature filter, which removes object-irrelevant confounders in a cyclical manner to extract the object-related causal features. Subsequently, to further enhance the causal feature invariance, we propose to train with augmented causal features by combining them with randomly-sampled styles from the object-irrelevant feature distribution boundary. As a result, our model bridges the gap between one seen domain and multiple unseen ones by capturing the invariant causal features, which largely enhances the model's generalization ability in SDG. In experiments, our method can be plugged into multiple state-of-the-art methods, and the significant performance improvements on multiple datasets demonstrate the superiority of our method. In particular, on the PACS dataset, our method achieves an accuracy improvement of 4.7%. Muli Yang, Aming Wu, Cheng Deng 0002 |
IEEE Trans. Image Process. | 2 |
| 2025 | Memory-Enhanced Confidence Calibration for Class-Incremental Unsupervised Domain AdaptationabstractIn this paper, we focus on Class-Incremental Unsupervised Domain Adaptation (CI-UDA), where the labeled source domain already includes all classes, and the classes in the unlabeled target domain emerge sequentially over time. This task involves addressing two main challenges. The first is the domain gap between the labeled source data and the unlabeled target data, which leads to weak generalization performance. The second is the inconsistency between the source and target category spaces at each time step, which causes catastrophic forgetting during the testing stage. Previous methods focus solely on the alignment of similar samples from different domains, which overlooks the underlying causes of the domain gap/class distribution difference. To tackle the issue, we rethink this task from a causal perspective for the first time. We first build a structural causal graph to describe the CI-UDA problem. Based on the causal graph, we present Memory-Enhanced Confidence Calibration (MECC), which aims to improve confidence in the predicted results. In particular, we argue that the domain discrepancy caused by the different styles is prone to make the model produce less confident predictions and thus weakens the generalization and continual learning abilities. To this end, we first explore using the gram matrix to generate source-style target data, which is combined with the original data to jointly train the model and thereby reduce the domain-shift impact. Second, we utilize the model of the previous time step to select corresponding samples that are used to build a memory bank, which is instrumental in alleviating catastrophic forgetting. Extensive experimental results on multiple datasets demonstrate the superiority of our method. Jiaping Yu, Muli Yang, Aming Wu, Cheng Deng 0002 |
IEEE Trans. Multim. | 2 |
| 2024 | LLM Knows Body Language, Too: Translating Speech Voices into Human GesturesabstractIn response to the escalating demand for digital human representations, progress has been made in the generation of realistic human gestures from given speeches.Despite the remarkable achievements of recent research, the generation process frequently includes unintended, meaningless, or non-realistic gestures.To address this challenge, we propose a gesture translation paradigm, GesTran, which leverages large language models (LLMs) to deepen the understanding of the connection between speech and gesture and sequentially generates human gestures by interpreting gestures as a unique form of body language.The primary stage of the proposed framework employs a transformer-based auto-encoder network to encode human gestures into discrete symbols.Following this, the subsequent stage utilizes a pre-trained LLM to decipher the relationship between speech and gesture, translating the speech into gesture by interpreting the gesture as unique language tokens within the LLM.Our method has demonstrated state-of-the-art performance improvement through extensive and impartial experiments conducted on public TED and TED-Expressive datasets. Guangtao Lyu, Jiexi Yan, Muli Yang, Cheng Deng 0002 |
ACL (1) | 4 |
| 2024 | MLGPnet: Multi-granularity neural network for 3D shape recognition using pyramid data
Zekun Li 0004, Seah Hock Soon, Baolong Guo 0001, Muli Yang |
Comput. Vis. Image Underst. | 4 |
| 2024 | Rethinking Noise Sampling in Class-Imbalanced Diffusion ModelsabstractIn the practical application of image generation, dealing with long-tailed data distributions is a common challenge for diffusion-based generative models. To tackle this issue, we investigate the head-class accumulation effect in diffusion models' latent space, particularly focusing on its correlation to the noise sampling strategy. Our experimental analysis indicates that employing a consistent sampling distribution for the noise prior across all classes leads to a significant bias towards head classes in the noise sampling distribution, which results in poor quality and diversity of the generated images. Motivated by this observation, we propose a novel sampling strategy named Bias-aware Prior Adjusting (BPA) to debias diffusion models in the class-imbalanced scenario. With BPA, each class is automatically assigned an adaptive noise sampling distribution prior during training, effectively mitigating the influence of class imbalance on the generation process. Extensive experiments on several benchmarks demonstrate that images generated using our proposed BPA showcase elevated diversity and superior quality. Jiexi Yan, Muli Yang, Cheng Deng 0002 |
IEEE Trans. Image Process. | 3 |
| 2023 | Bootstrap Your Own Prior: Towards Distribution-Agnostic Novel Class DiscoveryabstractNovel Class Discovery (NCD) aims to discover unknown classes without any annotation, by exploiting the transferable knowledge already learned from a base set of known classes. Existing works hold an impractical assumption that the novel class distribution prior is uniform, yet neglect the imbalanced nature of real-world data. In this paper, we relax this assumption by proposing a new challenging task: distribution-agnostic NCD, which allows data drawn from arbitrary unknown class distributions and thus renders existing methods useless or even harmful. We tackle this challenge by proposing a new method, dubbed “Boot-strapping Your Own Prior (BYOP)”, which iteratively estimates the class prior based on the model prediction it-self. At each iteration, we devise a dynamic temperature technique that better estimates the class prior by encouraging sharper predictions for less-confident samples. Thus, BYOP obtains more accurate pseudo-labels for the novel samples, which are beneficial for the next training iteration. Extensive experiments show that existing methods suffer from imbalanced class distributions, while BYOp11Code: https://github.com/muliyangm/BYOP. out-performs them by clear margins, demonstrating its effectiveness across various distribution scenarios. Muli Yang, Liancheng Wang, Cheng Deng 0002, Hanwang Zhang |
CVPR | 1 |
| 2023 | Hierarchical Prompt Learning for Compositional Zero-Shot RecognitionabstractCompositional Zero-Shot Learning (CZSL) aims to imitate the powerful generalization ability of human beings to recognize novel compositions of known primitive concepts that correspond to a state and an object, e.g., purple apple. To fully capture the intra- and inter-class correlations between compositional concepts, in this paper, we propose to learn them in a hierarchical manner. Specifically, we set up three hierarchical embedding spaces that respectively model the states, the objects, and their compositions, which serve as three “experts” that can be combined in inference for more accurate predictions. We achieve this based on the recent success of large-scale pretrained vision-language models, e.g., CLIP, which provides a strong initial knowledge of image-text relationships. To better adapt this knowledge to CZSL, we propose to learn three hierarchical prompts by explicitly fixing the unrelated word tokens in the three embedding spaces. Despite its simplicity, our proposed method consistently yields superior performance over current state-of-the-art approaches on three widely-used CZSL benchmarks. Henan Wang, Muli Yang, Cheng Deng 0002 |
IJCAI | 2 |
| 2023 | A Decomposable Causal View of Compositional Zero-Shot LearningabstractComposing and recognizing novel concepts that are combinations of known concepts,i.e., compositional generalization, is one of the greatest power of human intelligence. With the development of artificial intelligence, it becomes increasingly appealing to build a vision system that can generalize to unknown compositions based on restricted known knowledge, which has so far remained a great challenge to our community. In fact, machines can be easily misled by superficial correlations in the data, disregarding the causal patterns that are crucial to generalization. In this paper, we rethink compositional generalization with a causal perspective, upon the context of Compositional Zero-Shot Learning (CZSL). We develop a simple yet strong approach based on our novelDecomposableCausal view (dubbed “DeCa”), by approximating the causal effect with the combination of three easy-to-learn components. Our proposedDeCa11Code is available onhttps://github.com/muliyangm/DeCa.is evaluated on two challenging CZSL benchmarks by recognizing unknown compositions of known concepts. Despite being simple in the design, our approach achieves consistent improvements over state-of-the-art baselines, demonstrating its superiority towards the goal of compositional generalization. Muli Yang, Aming Wu, Cheng Deng 0002 |
IEEE Trans. Multim. | 1 |
| 2023 | Adaptive Bias-Aware Feature Generation for Generalized Zero-Shot LearningabstractZero-Shot Learning (ZSL) aims to recognize unseen classes that never appear during training. Recently, generative adversarial networks (GANs) have been introduced to convert ZSL into a supervised learning problem by synthesizing unseen visual features. However, since unseen classes are never experienced for the generator during training, the synthesized unseen visual features often become heavily biased towards seen classes, or sometimes there is even no meaningful class that can be assigned to them. This is known as thebias problem. In this paper, we propose a novel method, namely Adaptive Bias-Aware GAN (ABA-GAN), to alleviate generating biased visual features. For this purpose, we build a semantic adversarial network to regularize the feature generator. Specifically, an adaptive adversarial loss is proposed to constrain the feature distributions, which avoids the generation of meaningless visual features. Meanwhile, a domain divider is presented to explicitly distinguish synthesized visual features between seen and unseen domains, such that the bias towards seen classes can be alleviated. Moreover, we propose a novel metric named bias score (BS) to explicitly quantify the degree of the strong bias. Extensive experiments on four widely used benchmark datasets demonstrate that our proposed method outperforms the state-of-the-art approaches under both ZSL and GZSL protocols. Yanhua Yang, Xiaozhe Zhang, Muli Yang, Cheng Deng 0002 |
IEEE Trans. Multim. | 3 |
| 2022 | Siamese Contrastive Embedding Network for Compositional Zero-Shot LearningabstractCompositional Zero-Shot Learning (CZSL) aims to recognize unseen compositions formed from seen state and object during training. Since the same state may be various in the visual appearance while entangled with different objects, CZSL is still a challenging task. Some methods recognize state and object with two trained classifiers, ignoring the impact of the interaction between object and state; the other methods try to learn the joint representation of the state-object compositions, leading to the domain gap between seen and unseen composition sets. In this paper, we propose a novel Siamese Contrastive Embedding Network (SCEN)11Code: https://github.com/XDUxyLi/SCEN-master for unseen composition recognition. Considering the entanglement between state and object, we embed the visual feature into a Siamese Contrastive Space to capture prototypes of them separately, alleviating the interaction between state and object. In addition, we design a State Transition Module (STM) to increase the diversity of training compositions, improving the robustness of the recognition model. Extensive experiments indicate that our method significantly outperforms the state-of-the-art approaches on three challenging benchmark datasets, including the recent proposed C-QGA dataset. Xu Yang 0019, Cheng Deng 0002, Muli Yang |
CVPR | 5 |
| 2022 | Divide and Conquer: Compositional Experts for Generalized Novel Class DiscoveryabstractIn response to the explosively-increasing requirement of annotated data, Novel Class Discovery (NCD) has emerged as a promising alternative to automatically recognize unknown classes without any annotation. To this end, a model makes use of a base set to learn basic semantic discriminability that can be transferred to recognize novel classes. Most existing works handle the base and novel sets using separate objectives within a two-stage training paradigm. Despite showing competitive performance on novel classes, they fail to generalize to recognizing samples from both base and novel sets. In this paper, we focus on this generalized setting of NCD (GNCD), and propose to divide and conquer it with two groups of Compositional Experts (ComEx). Each group of experts is designed to characterize the whole dataset in a comprehensive yet complementary fashion. With their union, we can solve GNCD in an efficient end-to-end manner. We further look into the draw-back in current NCD methods, and propose to strengthen ComEx with global-to-local and local-to-local regularization. ComEx11Code: https://github.com/muliyangm/ComEx. is evaluated on four popular benchmarks, showing clear superiority towards the goal of GNCD. Muli Yang, Yuehua Zhu, Jiaping Yu, Aming Wu, Cheng Deng 0002 |
CVPR | 1 |
| 2022 | Progressive Self-Attention Network with Unsymmetrical Positional Encoding for Sequential RecommendationabstractIn real-world recommendation systems, the preferences of users are often affected by long-term constant interests and short-term temporal needs. The recently proposed Transformer-based models have proved superior in the sequential recommendation, modeling temporal dynamics globally via the remarkable self-attention mechanism. However, all equivalent item-item interactions in original self-attention are cumbersome, failing to capture the drifting of users' local preferences, which contain abundant short-term patterns. In this paper, we propose a novel interpretable convolutional self-attention, which efficiently captures both short- and long-term patterns with a progressive attention distribution. Specifically, a down-sampling convolution module is proposed to segment the overall long behavior sequence into a series of local subsequences. Accordingly, the segments are interacted with each item in the self-attention layer to produce locality-aware contextual representations, during which the quadratic complexity in original self-attention is reduced to nearly linear complexity. Moreover, to further enhance the robust feature learning in the context of Transformers, an unsymmetrical positional encoding strategy is carefully designed. Extensive experiments are carried out on real-world datasets, \eg ML-1M, Amazon Books, and Yelp, indicating that the proposed method outperforms the state-of-the-art methods w.r.t. both effectiveness and efficiency. Yuehua Zhu, Shaohua Jiang, Muli Yang, Yanhua Yang, Leon Wenliang Zhong |
SIGIR | 4 |
| 2020 | Learning Unseen Concepts via Hierarchical Decomposition and CompositionabstractComposing and recognizing new concepts from known sub-concepts has been a fundamental and challenging vision task, mainly due to 1) the diversity of sub-concepts and 2) the intricate contextuality between sub-concepts and their corresponding visual features. However, most of the current methods simply treat the contextuality as rigid semantic relationships and fail to capture fine-grained contextual correlations. We propose to learn unseen concepts in a hierarchical decomposition-and-composition manner. Considering the diversity of sub-concepts, our method decomposes each seen image into visual elements according to its labels, and learns corresponding sub-concepts in their individual subspaces. To model intricate contextuality between sub-concepts and their visual features, compositions are generated from these subspaces in three hierarchical forms, and the composed concepts are learned in a unified composition space. To further refine the captured contextual relationships, adaptively semi-positive concepts are defined and then learned with pseudo supervision exploited from the generated compositions. We validate the proposed approach on two challenging benchmarks, and demonstrate its superiority over state-of-the-art approaches. Muli Yang, Cheng Deng 0002, Junchi Yan, Xianglong Liu 0001, Dacheng Tao |
CVPR | 1 |
| 2020 | Progressive Domain-Independent Feature Decomposition Network for Zero-Shot Sketch-Based Image RetrievalabstractZero-Shot Sketch-Based Image Retrieval (ZS-SBIR) is a specific cross-modal retrieval task for searching natural images given free-hand sketches under the zero-shot scenario. Most existing methods solve this problem by simultaneously projecting visual features and semantic supervision into a low-dimensional common space for efficient retrieval. However, such low-dimensional projection destroys the completeness of semantic knowledge in original semantic space, so that it is unable to transfer useful knowledge well when learning semantic features from different modalities. Moreover, the domain information and semantic information are entangled in visual features, which is not conducive for cross-modal matching since it will hinder the reduction of domain gap between sketch and image. In this paper, we propose a Progressive Domain-independent Feature Decomposition (PDFD) network for ZS-SBIR. Specifically, with the supervision of original semantic knowledge, PDFD decomposes visual features into domain features and semantic ones, and then the semantic features are projected into common space as retrieval features for ZS-SBIR. The progressive projection strategy maintains strong semantic supervision. Besides, to guarantee the retrieval features to capture clean and complete semantic information, the cross-reconstruction loss is introduced to encourage that any combinations of retrieval features and domain features can reconstruct the visual features. Extensive experiments demonstrate the superiority of our PDFD over state-of-the-art competitors. Xinxun Xu, Muli Yang, Yanhua Yang, Hao Wang 0062 |
IJCAI | 2 |
| 2020 | Fewer is More: A Deep Graph Metric Learning Perspective Using Fewer ProxiesabstractDeep metric learning plays a key role in various machine learning tasks. Most of the previous works have been confined to sampling from a mini-batch, which cannot precisely characterize the global geometry of the embedding space. Although researchers have developed proxy- and classification-based methods to tackle the sampling issue, those methods inevitably incur a redundant computational cost. In this paper, we propose a novel Proxy-based deep Graph Metric Learning (ProxyGML) approach from the perspective of graph classification, which uses fewer proxies yet achieves better comprehensive performance. Specifically, multiple global proxies are leveraged to collectively approximate the original data points for each class. To efficiently capture local neighbor relationships, a small number of such proxies are adaptively selected to construct similarity subgraphs between these proxies and each data point. Further, we design a novel reverse label propagation algorithm, by which the neighbor relationships are adjusted according to ground-truth labels, so that a discriminative metric space can be learned during the process of subgraph classification. Extensive experiments carried out on widely-used CUB-200-2011, Cars196, and Stanford Online Products datasets demonstrate the superiority of the proposed ProxyGML over the state-of-the-art methods in terms of both effectiveness and efficiency. The source code is publicly available at \url{https://github.com/YuehuaZhu/ProxyGML}. Yuehua Zhu, Muli Yang, Cheng Deng 0002, Wei Liu 0005 |
NeurIPS | 2 |
| 2020 | Progressive Cross-Modal Semantic Network for Zero-Shot Sketch-Based Image RetrievalabstractZero-shot sketch-based image retrieval (ZS-SBIR) is a specific cross-modal retrieval task that involves searching natural images through the use of free-hand sketches under the zero-shot scenario. Most previous methods project the sketch and image features into a low-dimensional common space for efficient retrieval, and meantime align the projected features to their semantic features (e.g., category-level word vectors) in order to transfer knowledge from seen to unseen classes. However, the projection and alignment are always coupled; as a result, there is a lack of alignment that consequently leads to unsatisfactory zero-shot retrieval performance. To address this issue, we propose a novel progressive cross-modal semantic network. More specifically, it first explicitly aligns the sketch and image features to semantic features, then projects the aligned features to a common space for subsequent retrieval. We further employ cross-reconstruction loss to encourage the aligned features to capture complete knowledge about the two modalities, along with multi-modal Euclidean loss that guarantees similarity between the retrieval features from a sketch-image pair. Extensive experiments conducted on two popular large-scale datasets demonstrate that our proposed approach outperforms state-of-the-art competitors to a remarkable extent: by more than 3% on the Sketchy dataset and about 6% on the TU-Berlin dataset in terms of retrieval accuracy. Cheng Deng 0002, Xinxun Xu, Hao Wang 0062, Muli Yang, Dacheng Tao |
IEEE Trans. Image Process. | 4 |
| 2019 | Adversarial Fine-Grained Composition Learning for Unseen Attribute-Object RecognitionabstractRecognizing unseen attribute-object pairs never appearing in the training data is a challenging task, since an object often refers to a specific entity while an attribute is an abstract semantic description. Besides, attributes are highly correlated to objects, i.e., an attribute tends to describe different visual features of various objects. Existing methods mainly employ two classifiers to recognize attribute and object separately, or simply simulate the composition of attribute and object, which ignore the inherent discrepancy and correlation between them. In this paper, we propose a novel adversarial fine-grained composition learning model for unseen attribute-object pair recognition. Considering their inherent discrepancy, we leverage multi-scale feature integration to capture discriminative fine-grained features from a given image. Besides, we devise a quintuplet loss to depict more accurate correlations between attributes and objects. Adversarial learning is employed to model the discrepancy and correlations among attributes and objects. Extensive experiments on two challenging benchmarks indicate that our method consistently outperforms state-of-the-art competitors by a large margin. Muli Yang, Hao Wang 0062, Cheng Deng 0002, Xianglong Liu 0001 |
ICCV | 2 |
| 2019 | Adaptive-weighting discriminative regression for multi-view classification
Muli Yang, Cheng Deng 0002, Feiping Nie 0001 |
Pattern Recognit. | 1 |