EDBT 2026 Demo / reviewers in the wild / expert
Man Liu 0003
dblp:63/8087-3
· DBLP profile ↗
10ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0002-8062-5566ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Harnessing Group-Oriented Consistency Constraints for Semi-Supervised Semantic Segmentation in CdZnTe Semiconductors
Man Liu 0003, Huihui Bai 0001, Anhong Wang, Yunchao Wei, Yao Zhao 0001 |
IEEE Trans. Image Process. | 3 |
| 2026 | Deep Multi-View Clustering With Intra-View Similarity and Cross-View Correlation LearningabstractDeep multi-view clustering (MVC) has gained widespread attention as it can effectively mine consistent information from multiple views and improve clustering performance. However, view bias often exists between views (i.e., the quality differences between views). Treating all views equally inevitably destroys structural information when simply concatenating or summing the embedded representation of multiple views. To alleviate this issue, we propose a deep multi-view clustering with intra-view similarity and cross-view correlation learning (MISCC), facilitating the intra-view discriminability and inter-view complementarity. Specifically, we utilize the intra-view inherent structure information to dynamically identify semantically similar samples within each view. By aggregating their embedding representations, fine-grained structures are enhanced to boost intra-cluster compactness and inter-cluster separation. Then, we construct a cross-view correlation learning module to align semantically related views while preserving the distinctive features of irrelevant views. Based on them, a centralized clustering alignment strategy is proposed to align the similarity distribution and clustering structure between each view and the unified view, balancing the diverse information among multiple views. By jointly training these modules, the unified representation is optimized to capture more discriminative information from multiple views. Extensive experiments conducted on eleven multi-view datasets demonstrate that MISCC outperforms the state-of-the-art clustering methods. Pengyuan Li 0013, Dongxia Chang, Yiming Wang 0007, Man Liu 0003, Zisen Kong, Linhua Kong, Yao Zhao 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | Attend and Enrich: Enhanced Visual Prompt for Zero-Shot LearningabstractZero-shot learning (ZSL) endeavors to transfer knowledge from the seen categories to recognize unseen categories, which mostly relies on the semantic-visual interactions between image and attribute tokens. Recently, the prompt learning has emerged in ZSL and demonstrated significant potential as it allows the zero-shot transfer of diverse visual concepts to downstream tasks. However, current methods explore the fixed adaptation of the learnable prompt on the seen domains, which make them over-emphasize the primary visual features observed during training, limiting their generalization capabilities to the unseen domains. In this work, we propose AENet, which endows semantic information into the visual prompt to distill semantic-enhanced prompt for visual representation enrichment, enabling effective knowledge transfer for ZSL. AENet comprises two key steps: 1) exploring the concept-harmonized tokens for the visual and attribute modalities, grounded on the modal-sharing token that represents consistent visual-semantic concepts; and 2) yielding the semantic-enhanced prompt via the visual residual refinement unit with attribute consistency supervision. It is further integrated with primary visual features to attend to semantic-related information for visual enhancement, thus strengthening transferable ability. Experimental results on three benchmarks show that our AENet outperforms existing state-of-the-art ZSL methods. Man Liu 0003, Huihui Bai 0001, Feng Li 0037, Chunjie Zhang 0001, Yunchao Wei, Tat-Seng Chua, Yao Zhao 0001 |
AAAI | 1 |
| 2025 | AEMVC: Mitigate Imbalanced Embedding Space in Multi-view ClusteringabstractMulti-view clustering (MVC) has gained extensive attention for its capacity to handle heterogeneous data. However, current autoencoder-based MVC methods suffer from a limitation: embedding space exhibits severe imbalances in the efficacy of feature direction, creating a long-tailed singular value distribution where few directions dominate. To mitigate this, we introduce a novel Activate-Then-Eliminate Strategy for Multi-View Clustering (AEMVC), inspired by the observation that balanced feature directions can facilitate enhancing discrimination of learned representations. AEMVC dynamically adjusts the contributions of different feature directions through two keys: a Feature Activation Module that narrows singular value discrepancies to prevent dominant directions from controlling clustering decisions, and an Inter-view Mutual Supervision strategy that filters redundant information by adaptively determining view-specific thresholds based on cross-view consistency. By activating more feature directions and eliminating each view's adverse factors, AEMVC achieves more balanced and discriminative embedding representations. Extensive experiments on seven multi-view benchmarks validate AEMVC's effectiveness, demonstrating substantial improvements over state-of-the-art methods. Pengyuan Li 0013, Man Liu 0003, Dongxia Chang, Yiming Wang 0007, Zisen Kong, Yao Zhao 0001 |
ACM Multimedia | 2 |
| 2025 | PSVMA+: Exploring Multi-Granularity Semantic-Visual Adaption for Generalized Zero-Shot LearningabstractGeneralized zero-shot learning (GZSL) endeavors to identify the unseen categories using knowledge from the seen domain, necessitating the intrinsic interactions between the visual features and attribute semantic features. However, GZSL suffers from insufficient visual-semantic correspondences due to the attribute diversity and instance diversity. Attribute diversity refers to varying semantic granularity in attribute descriptions, ranging from low-level (specific, directly observable) to high-level (abstract, highly generic) characteristics. This diversity challenges the collection of adequate visual cues for attributes under a uni-granularity. Additionally, diverse visual instances corresponding to the same sharing attributes introduce semantic ambiguity, leading to vague visual patterns. To tackle these problems, we propose a multi-granularity progressive semantic-visual mutual adaption (PSVMA+) network, where sufficient visual elements across granularity levels can be gathered to remedy the granularity inconsistency. PSVMA+ explores semantic-visual interactions at different granularity levels, enabling awareness of multi-granularity in both visual and semantic elements. At each granularity level, the dual semantic-visual transformer module (DSVTM) recasts the sharing attributes into instance-centric attributes and aggregates the semantic-related visual regions, thereby learning unambiguous visual features to accommodate various instances. Given the diverse contributions of different granularities, PSVMA+ employs selective cross-granularity learning to leverage knowledge from reliable granularities and adaptively fuses multi-granularity features for comprehensive representations. Experimental results demonstrate that PSVMA+ consistently outperforms state-of-the-art methods. Man Liu 0003, Huihui Bai 0001, Feng Li 0037, Chunjie Zhang 0001, Yunchao Wei, Meng Wang 0001, Tat-Seng Chua, Yao Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Exploring Resolution Fields for Scalable Image Compression With Uncertainty GuidanceabstractRecently, there are significant advancements in learning-based image compression methods surpassing traditional coding standards. Most of them prioritize achieving the best rate-distortion performance for a particular compression rate, which limits their flexibility and adaptability in various applications with complex and varying constraints. In this work, we explore the potential of resolution fields in scalable image compression and propose the reciprocal pyramid network (RPN) that fulfills the need for more adaptable and versatile compression. Specifically, RPN first builds a compression pyramid and generates the resolution fields at different levels in a top-down manner. The key design lies in the cross-resolution context mining module between adjacent levels, which performs feature enriching and distillation to mine meaningful contextualized information and remove unnecessary redundancy, producing informative resolution fields as residual priors. The scalability is achieved by progressive bitstream reusing and resolution field incorporation varying at different levels. Furthermore, between adjacent compression levels, we explicitly quantify the aleatoric uncertainty from the bottom decoded representations and develop an uncertainty-guided loss to update the upper-level compression parameters, forming a reverse pyramid process that enforces the network to focus on the textured pixels with high variance for more reliable and accurate reconstruction. Combining resolution field exploration and uncertainty guidance in a pyramid manner, RPN can effectively achieve spatial and quality scalable image compression. Experiments show the superiority of RPN against existing classical and deep learning-based scalable codecs. Code will be available athttps://github.com/JGIroro/RPNSIC. Dongyi Zhang, Feng Li 0037, Man Liu 0003, Runmin Cong, Huihui Bai 0001, Meng Wang 0001, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Progressive Complementary Knowledge Aggregation for CdZnTe Defect SegmentationabstractAutomatic quality inspection of industrial products is an indispensable part of modern manufacturing. Cadmium zinc telluride (CdZnTe) crystal is an important industrial raw material, but the special photosensitive properties of CdZnTe make it show different defect boundaries under different lighting angles, which poses challenges for quality inspection. In this article, we propose progressive complementary knowledge aggregation (PCKA) for CdZnTe defect segmentation, which is model-agnostic. First, the 12 images of CdZnTe crystal with different lighting angles are fed into the preliminary aggregation net to aggregate unique pixel-level clues. Second, we use a latent aggregation net to acquire the feature-level complementary clues under the guidance of the pixel-level clues within latent space. Such a learning paradigm is an effective solution for the special photosensitive properties of CdZnTe crystal. Extensive experiments on self-collected dataset demonstrate the effectiveness and efficiency of our PCKA compared with other solutions. Feng Li 0037, Man Liu 0003, Huihui Bai 0001, Yunchao Wei, Anhong Wang, Shijie Ma, Yao Zhao 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2024 | Part-Object Progressive Refinement Network for Zero-Shot LearningabstractZero-shot learning (ZSL) recognizes unseen images by sharing semantic knowledge transferred from seen images, encouraging the investigation of associations between semantic and visual information. Prior works have been devoted to the alignment of global visual features with semantic information, i.e., attribute vectors, or further mining the local part regions related to each attribute and then simply concatenating them for category decisions. Although effective, these works ignore intrinsic interactions between local parts and the whole object, which enables a more discriminative and representative knowledge transfer for ZSL. In this paper, we propose a Part-Object Progressive Refinement Network (POPRNet), where discriminative and transferable semantics are progressively refined by the cooperation between parts and the whole object. Specifically, POPRNet incorporates discriminative part semantics and object-centric semantics guided by semantic intensity to improve cross-domain transferability. To achieve part-object learning, a semantic-augment transformer (SaT) is proposed to model the part-object relation at the part-level via an encoder and at the object-level via a decoder, generating a comprehensive semantic representation to boost discriminability and transferability. By introducing the prototype updating module embedded with the prototype selection layers, the discriminative ability of the updated category prototype is enhanced to further improve the recognition performance of ZSL. Extensive experiments are conducted to demonstrate the superiority and competitiveness of our proposed POPRNet method on three public benchmark datasets. The code is available at https://github.com/ManLiuCoder/POPRNet. Man Liu 0003, Chunjie Zhang 0001, Huihui Bai 0001, Yao Zhao 0001 |
IEEE Trans. Image Process. | 1 |
| 2023 | Progressive Semantic-Visual Mutual Adaption for Generalized Zero-Shot LearningabstractGeneralized Zero-Shot Learning (GZSL) identifies unseen categories by knowledge transferred from the seen domain, relying on the intrinsic interactions between visual and semantic information. Prior works mainly localize regions corresponding to the sharing attributes. When various visual appearances correspond to the same attribute, the sharing attributes inevitably introduce semantic ambiguity, hampering the exploration of accurate semantic-visual interactions. In this paper, we deploy the dual semantic-visual transformer module (DSVTM) to progressively model the correspondences between attribute prototypes and visual features, constituting a progressive semantic-visual mutual adaption (PSVMA) network for semantic disambiguation and knowledge transferability improvement. Specifically, DSVTM devises an instance-motivated semantic encoder that learns instance-centric prototypes to adapt to different images, enabling the recast of the unmatched semantic-visual pair into the matched one. Then, a semantic-motivated instance decoder strengthens accurate cross-domain interactions between the matched pair for semantic-related instance adaption, en-couraging the generation of unambiguous visual representations. Moreover, to mitigate the bias towards seen classes in GZSL, a debiasing loss is proposed to pursue response consistency between seen and unseen predictions. The PSVMA consistently yields superior performances against other state-of-the-art methods. Code will be available at: https://github.com/ManLiuCoder/PSVMA. Man Liu 0003, Feng Li 0037, Chunjie Zhang 0001, Yunchao Wei, Huihui Bai 0001, Yao Zhao 0001 |
CVPR | 1 |
| 2022 | Cross-Part Learning for Fine-Grained Image ClassificationabstractRecent techniques have achieved remarkable improvements depended on mining subtle yet distinctive features for fine-grained visual classification (FGVC). While prior works directly combine discriminative features extracted from different parts, we argue that the potential interactions between different parts and their abilities to category predictions should be taken into consideration, which enables significant parts to contribute more to the decision of the sub-category. To this end, we present a Cross-Part Convolutional Neural Network (CP-CNN) in a weakly supervised manner to explore cross-learning among multi-regional features. Specifically, the context transformer is implemented to encourage joint feature learning across different parts under the guidance of a navigator. The part with the highest confidence is regarded as a navigator to deliver distinguishing characteristics to the others with lower confidence while the complementary information is retained. To locate discriminative but subtle parts precisely, a part proposal generator (PPG) is designed with the feature enhancement blocks, through which complex scale variations caused by the viewpoint diversity can be effectively alleviated. Extensive experiments on three benchmark datasets demonstrate that our proposed method consistently outperforms existing state-of-the-art methods. Man Liu 0003, Chunjie Zhang 0001, Huihui Bai 0001, Riquan Zhang, Yao Zhao 0001 |
IEEE Trans. Image Process. | 1 |