VLDB 2026 Research / reviewers in the wild / expert
Man Zhang 0005
dblp:49/5096-5
· DBLP profile ↗
35ranked-venue papers
4as first author
23since 2021 · last 2026
0000-0003-3043-2122ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 27 · 3 first-author · 19 since 2021Artificial intelligence and machine learning · 21 · 3 first-author · 12 since 2021Security and privacy · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive InteractionabstractGenerating responsive listener head dynamics with nuanced emotions and expressive reactions is crucial for dialogue modeling in various virtual avatar animations. Previous studies mainly focus on the direct short-term production of listener behavior. They overlook the fine-grained control over motion variations and emotional intensity, especially in long-sequence modeling. Moreover, the lack of long-term and large-scale paired speaker-listener corpora incorporating head dynamics and fine-grained multi-modality annotations limits the application of dialogue modeling. Therefore, we first newly collect a large-scale multi-turn dataset of 3D dyadic conversation containing more than 1.4M valid frames for multi-modal responsive interaction, dubbed ListenerX. Additionally, we propose VividListener, a novel framework enabling fine-grained, expressive, and controllable listener dynamics modeling. This framework leverages multi-modal conditions as guiding principles for fostering coherent interactions between speakers and listeners. Specifically, we design the Responsive Interaction Module (RIM) to adaptively represent the multi-modal interactive embeddings. RIM ensures the listener dynamics achieve fine-grained semantic coordination with textual descriptions and adjustments, while preserving expressive reaction with speaker behavior. Meanwhile, we propose the Emotional Intensity Tags (EIT) for emotion intensity editing with multi-modal information integration, applying to both text descriptions and listener motion amplitude. Extensive experiments conducted on our newly collected ListenerX dataset demonstrate that VividListener achieves state-of-the-art performance, realizing expressive and controllable listener dynamics. Xingqun Qi, Bingkun Yang, Weile Chen, Zezhao Tian, Muyi Sun, Man Zhang 0005, Zhenan Sun |
AAAI | 8 |
| 2026 | FurniScene: A Large-scale 3D Room Dataset with Intricate Furnishing Scenes
Yuxi Wang 0001, Junran Peng, Genghao Zhang, Chuanchen Luo, Shibiao Xu, Man Zhang 0005, Zhaoxiang Zhang 0001 |
Int. J. Comput. Vis. | 6 |
| 2026 | Rating-aware argument generation for movie reviews with multimodal large language models and a new dataset
Wenjie Hua, Quan Fang, Muyi Sun, Shibiao Xu, Man Zhang 0005 |
Multim. Syst. | 5 |
| 2025 | OpenAnimals: Revisiting Person Re-Identification for Animals Towards Better GeneralizationabstractThis paper addresses the challenge of animal re-identification, an emerging field that shares similarities with person re-identification but presents unique complexities due to the diverse species, environments and poses. To facilitate research in this domain, we introduce OpenAnimals, a flexible and extensible codebase designed specifically for animal re-identification. We conduct a comprehensive study by revisiting several state-of-the-art person re-identification methods, including BoT, AGW, SBS, and MGN, and evaluate their effectiveness on animal re-identification benchmarks such as HyenaID, LeopardID, SeaTurtleID, and WhaleSharkID. Our findings reveal that while some techniques generalize well, many do not, underscoring the significant differences between the two tasks. To bridge this gap, we propose ARBase, a strong \textbf{Base} model tailored for \textbf{A}nimal \textbf{R}e-identification, which incorporates insights from extensive experiments and introduces simple yet effective animal-oriented designs. Experiments demonstrate that ARBase consistently outperforms existing baselines, achieving state-of-the-art performance across various benchmarks. Saihui Hou, Panjian Huang, Zengbin Wang, Yuan Liu 0043, Man Zhang 0005, Yongzhen Huang |
ICCV | 6 |
| 2025 | Gait: Exploring X Modality for Generalized Gait Recognition
Zengbin Wang, Saihui Hou, Junjie Li 0002, Xu Liu 0008, Chunshui Cao, Yongzhen Huang, Siye Wang, Man Zhang 0005 |
ICCV | 8 |
| 2025 | DanceEditor: Towards Iterative Editable Music-Driven Dance Generation with Open-Vocabulary DescriptionsabstractGenerating coherent and diverse human dances from music signals has gained tremendous progress in animating virtual avatars. While existing methods support direct dance synthesis, they fail to recognize that enabling users to edit dance movements is far more practical in real-world choreography scenarios. Moreover, the lack of high-quality dance datasets incorporating iterative editing also limits addressing this challenge. To achieve this goal, we first construct DanceRemix, a large-scale multiturn editable dance dataset comprising the prompt featuring over 25.3 M dance frames and 84.5 K pairs. In addition, we propose a novel framework for iterative and editable dance generation coherently aligned with given music signals, namely DanceEditor. Considering the dance motion should be both musical rhythmic and enable iterative editing by user descriptions, our framework is built upon a prediction-then-editing paradigm unifying multimodal conditions. At the initial prediction stage, our framework improves the authority of generated results by directly modeling dance movements from tailored, aligned music. Moreover, at the subsequent iterative editing stages, we incorporate text descriptions as conditioning information to draw the editable results through a specifically designed Cross-modality Editing Module (CEM). Specifically, CEM adaptively integrates the initial prediction with music and text prompts as temporal motion cues to guide the synthesized sequences. Thereby, the results display music harmonics while preserving fine-grained semantic alignment with text descriptions. Extensive experiments demonstrate that our method outperforms the state-of-the-art models on our newly collected DanceRemix dataset. Code is available at https://lzvsdy.github.io/DanceEditor/. Xingqun Qi, Muyi Sun, Siye Wang, Man Zhang 0005, Sirui Han |
ICCV | 7 |
| 2025 | ReMeREC: Relation-aware and Multi-entity Referring Expression ComprehensionabstractReferring Expression Comprehension (REC) aims to localize specified entities or regions from the source image according to the given natural language descriptions. While existing methods enable single-entity localization, they overlook modeling the complex inter-entity relationship in more practical multi-entity scenes, which limits their ability to produce accurate and reliable results. Moreover, the lack of high-quality multi-entity datasets incorporating fine-grained and paired image-text-relation annotations also limits addressing this challenge. To achieve this task, we first manually construct a relation-aware multi-entity REC dataset with fine-grained relation and text annotations, namely ReMeX. Additionally, we propose ReMeREC, a novel framework that effectively integrates textual and visual cues to localize multiple entities while capturing their inter-relationship. Specifically, to mitigate the semantic ambiguity arising from the absence of explicit entity boundaries in the source natural language description, we introduce a novel Text-adaptive Multi-entity Perceptron (TMP). TMP dynamically infers both the quantity and span of entities from corresponding fine-grained text cues, thus deriving representations that preserve the unique characteristics of each entity. Meanwhile, we design the Entity Inter-relationship Reasoner (EIR) to enhance semantic distinctiveness relationship modeling, leading to a more profound perception of the global scene. Furthermore, to better capture the fine-grained linguistic prompts for delineating multiple entity boundaries and inter-relationship, we leverage LLMs to generate a small-scale textual dataset, dubbed EntityText, which serves as an effective auxiliary resource and further improves the textual understanding. Extensive experiments conducted on four benchmark datasets demonstrate the superior performance of our framework. Remarkably, ReMeREC achieves outstanding results in multi-entity grounding and complex relationship prediction, outperforming other counterparts by a large margin. Yizhi Hu, Zezhao Tian, Xingqun Qi, Bingkun Yang, Junhui Yin, Muyi Sun, Man Zhang 0005, Zhenan Sun |
ACM Multimedia | 8 |
| 2025 | Edge-Oriented Adversarial Attack for Deep Gait Recognition
Saihui Hou, Zengbin Wang, Man Zhang 0005, Chunshui Cao, Xu Liu 0008, Yongzhen Huang |
Int. J. Comput. Vis. | 3 |
| 2025 | Multi-Scale Semantic-Guidance Networks: Robust Blind Face Restoration Against Adversarial AttacksabstractImage processing networks are known to be vulnerable to adversarial examples, where adding carefully crafted adversarial perturbations to the inputs can mislead the model. This paper addresses the problem of robust blind face restoration (BFR) against adversarial attacks. BFR refers to recovering the HQ images from the LQ images, which suffer from diverse unknown degradation, such as noise, blur, artifact removal, low resolution, etc. Although existing BFR methods exhibit good performance, they experience significant degradation when subtle distortions and perturbations are introduced into the input images. This paper is the first to investigate, improve comprehensively, and evaluate BFR methods towards adversarial attacks. Project Gradient Descent (PGD) is employed to generate adversarial examples, and multiple types of attacks were used to thoroughly assess the robustness of various BFR methods across different objectives, regions, and levels. We evaluate the robustness of multiple BFR methods and analyze the advantages of their structures and modules towards adversarial attacks. Experimental results demonstrate that the method utilizing latent feature encoding and pre-trained discrete HQ codebook achieves better robustness than other methods, with the latter outperforming the former. Similarly, multi-scale semantic guidance information also exhibits superior performance in enhancing robustness. Therefore, we propose a powerful BFR method to mitigate this issue while maintaining better performance. Extensive experiments on three real-world datasets demonstrate our method’s state-of-the-art robustness in different scenarios. Zhenyuan Zhang 0001, Xingqun Qi, Zhenbo Song, Zhiqin Yang, Jianfeng Lu 0003, Muyi Sun, Man Zhang 0005, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2025 | Token Masking Transformer for Weakly Supervised Object LocalizationabstractWeakly supervised object localization (WSOL) is both a promising and challenging task that aims to achieve object localization exclusively through image category labels for supervision. Visual transformers have recently been applied to WSOL, demonstrating significant success through the exploitation of long-range feature dependencies in self-attention mechanisms. However, the transformer-based approach suffers from the same partial activation problem as the CNN-based approach due to the use of the classification task to train self-attention map, i.e., only a few discriminative regions are assigned high attention response and thus the localization map does not cover the whole object. To alleviate this problem, we propose a plug-and-play Token Masking Transformer (TMT) method to help transformer-based WSOL methods to obtain a more complete localization map by dynamic discriminative token masking. Specifically, a batch-wise discriminative token selection strategy is first introduced to flexibly determine the tokens to be masked in each image. Then, we design a token masking transformer block to perform token masking and inspire the network to mine more object-related tokens. Besides, we also design an intermediate token activation loss to further improve the performance of TMT by imposing constraints on intermediate tokens. Extensive experiments demonstrate that our TMT can substantially improve the performance of existing transformer-based methods without increasing the computational cost, and achieves state-of-the-art performance on two mainstream benchmarks. Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Man Zhang 0005, Xiaopeng Zhang 0001 |
IEEE Trans. Multim. | 6 |
| 2024 | QAGait: Revisit Gait Recognition from a Quality PerspectiveabstractGait recognition is a promising biometric method that aims to identify pedestrians from their unique walking patterns. Silhouette modality, renowned for its easy acquisition, simple structure, sparse representation, and convenient modeling, has been widely employed in controlled in-the-lab research. However, as gait recognition rapidly advances from in-the-lab to in-the-wild scenarios, various conditions raise significant challenges for silhouette modality, including 1) unidentifiable low-quality silhouettes (abnormal segmentation, severe occlusion, or even non-human shape), and 2) identifiable but challenging silhouettes (background noise, non-standard posture, slight occlusion). To address these challenges, we revisit gait recognition pipeline and approach gait recognition from a quality perspective, namely QAGait. Specifically, we propose a series of cost-effective quality assessment strategies, including Maxmial Connect Area and Template Match to eliminate background noises and unidentifiable silhouettes, Alignment strategy to handle non-standard postures. We also propose two quality-aware loss functions to integrate silhouette quality into optimization within the embedding space. Extensive experiments demonstrate our QAGait can guarantee both gait reliability and performance enhancement. Furthermore, our quality assessment strategies can seamlessly integrate with existing gait datasets, showcasing our superiority. Code is available at https://github.com/wzb-bupt/QAGait. Zengbin Wang, Saihui Hou, Man Zhang 0005, Xu Liu 0008, Chunshui Cao, Yongzhen Huang, Shibiao Xu |
AAAI | 3 |
| 2024 | Spectral Prompt Tuning: Unveiling Unseen Classes for Zero-Shot Semantic SegmentationabstractRecently, CLIP has found practical utility in the domain of pixel-level zero-shot segmentation tasks. The present landscape features two-stage methodologies beset by issues such as intricate pipelines and elevated computational costs. While current one-stage approaches alleviate these concerns and incorporate Visual Prompt Training (VPT) to uphold CLIP's generalization capacity, they still fall short in fully harnessing CLIP's potential for pixel-level unseen class demarcation and precise pixel predictions. To further stimulate CLIP's zero-shot dense prediction capability, we propose SPT-SEG, a one-stage approach that improves CLIP's adaptability from image to pixel. Specifically, we initially introduce Spectral Prompt Tuning (SPT), incorporating spectral prompts into the CLIP visual encoder's shallow layers to capture structural intricacies of images, thereby enhancing comprehension of unseen classes. Subsequently, we introduce the Spectral Guided Decoder (SGD), utilizing both high and low-frequency information to steer the network's spatial focus towards more prominent classification features, enabling precise pixel-level prediction outcomes. Through extensive experiments on two public datasets, we demonstrate the superiority of our method over state-of-the-art approaches, performing well across all classes and particularly excelling in handling unseen classes. Rongtao Xu, Changwei Wang 0001, Shibiao Xu, Li Guo 0004, Man Zhang 0005, Xiaopeng Zhang 0001 |
AAAI | 6 |
| 2024 | HardMo: A Large-Scale Hardcase Dataset for Motion CaptureabstractRecent years have witnessed rapid progress in monoc-ular human mesh recovery. Despite their impressive performance on public benchmarks, existing methods are vulnerable to unusual poses, which prevents them from deploying to challenging scenarios such as dance and martial arts. This issue is mainly attributed to the domain gap induced by the data scarcity in relevant cases. Most existing datasets are captured in constrained scenarios and lack samples of such complex movements. For this reason, we propose a data collection pipeline comprising automatic crawling, precise annotation, and hardcase mining. Based on this pipeline, we establish a large dataset in a short time. The dataset, named HardMo, contains 7M images along with precise annotations covering 15 categories of dance and 14 categories of martial arts. Empirically, we find that the prediction failure in dance and martial arts is mainly characterized by the misalignment of hand-wrist and foot-ankle. To dig deeper into the two hardcases, we leverage the proposed automatic pipeline to filter collected data and construct two subsets named HardMo-Hand and HardMo-Foot. Extensive experiments demonstrate the effectiveness of the annotation pipeline and the data-driven solution to failure cases. Specifically, after being trained on HardMo, HMR, an early pioneering method, can even outperform the current state of the art, 4DHumans, on our benchmarks. Dataset will be publicly available at https://ljqnb.github.io/HardMo.github.io. Jiaqi Liao, Chuanchen Luo, Yinuo Du, Yuxi Wang 0001, Xu-Cheng Yin, Man Zhang 0005, Zhaoxiang Zhang 0001, Junran Peng |
CVPR | 6 |
| 2024 | DefFusion: Deformable Multimodal Representation Fusion for 3D Semantic SegmentationabstractThe complementarity between camera and LiDAR data makes fusion methods a promising approach to improve 3D semantic segmentation performance. Recent transformer-based methods have also demonstrated superiority in segmentation. However, multimodal solutions incorporating transformers are underexplored and face two key inherent difficulties: over-attention and noise from different modal data. To overcome these challenges, we propose a Deformable Multimodal Representation Fusion (DefFusion) framework consisting mainly of a Deformable Representation Fusion Transformer and Dynamic Representation Augmentation Modules. The Deformable Representation Fusion Transformer introduces the deformable mechanism in multimodal fusion, avoiding over-attention and improving efficiency by adaptively modeling a 2D key/value set for a given 3D query, thus enabling multimodal fusion with higher flexibility. To enhance the 2D representation and 3D representation, the Dynamic Representation Enhancement Module is proposed to dynamically remove noise in the input representation via Dynamic Grouped Representation Generation and Dynamic Mask Generation. Extensive experiments validate that our model achieves the best 3D semantic segmentation performance on SemanticKITTI and NuScenes benchmarks. Rongtao Xu, Changwei Wang 0001, Duzhen Zhang, Man Zhang 0005, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001 |
ICRA | 4 |
| 2024 | Probabilistic Contrastive Learning for Domain Adaptation
Junjie Li 0002, Yixin Zhang 0007, Zilei Wang, Saihui Hou, Keyu Tu, Man Zhang 0005 |
IJCAI | 6 |
| 2024 | StableMoFusion: Towards Robust and Efficient Diffusion-based Motion Generation FrameworkabstractThanks to the powerful generative capacity of diffusion models, recent years have witnessed rapid progress in human motion generation. Existing diffusion-based methods employ disparate network architectures and training strategies. The effect of the design of each component is still unclear. In addition, the iterative denoising process consumes considerable computational overhead, which is prohibitive for real-time scenarios such as virtual characters and humanoid robots. For this reason, we first conduct a comprehensive investigation into network architectures, training strategies, and inference process. Based on the profound analysis, we tailor each component for efficient high-quality human motion generation. Despite the promising performance, the tailored model still suffers from foot skating which is an ubiquitous issue in diffusion-based solutions. To eliminate footskate, we identify foot-ground contact and correct foot motions along the denoising process. By organically combining these well-designed components together, we present StableMoFusion, a robust and efficient framework for human motion generation. Extensive experimental results show that our StableMoFusion performs favorably against current state-of-the-art methods. Chuanchen Luo, Yuxi Wang 0001, Shibiao Xu, Zhaoxiang Zhang 0001, Man Zhang 0005, Junran Peng |
ACM Multimedia | 7 |
| 2024 | MaterialSeg3D: Segmenting Dense Materials from 2D Priors for 3D AssetsabstractDriven by powerful image diffusion models, recent research has achieved the automatic creation of 3D objects from textual or visual guidance. By performing score distillation sampling (SDS) iteratively across different views, these methods succeed in lifting 2D generative prior to the 3D space. However, such a 2D generative image prior bakes the effect of illumination and shadow into the texture. As a result, material maps optimized by SDS inevitably involve spurious correlated components. The absence of precise material definition makes it infeasible to relight the generated assets reasonably in novel scenes, which limits their application in downstream scenarios. In contrast, humans can effortlessly circumvent this ambiguity by deducing the material of the object from its appearance and semantics. Motivated by this insight, we propose MaterialSeg3D, a 3D asset material generation framework to infer underlying material from the 2D semantic prior. Based on such a prior model, we devise a mechanism to parse material in 3D space. We maintain a UV stack, each map of which is unprojected from a specific viewpoint. After traversing all viewpoints, we fuse the stack through a weighted voting scheme and then employ region unification to ensure the coherence of the object parts. To fuel the learning of semantics prior, we collect a material dataset, named Materialized Individual Objects (MIO), which features abundant images, diverse categories, and accurate annotations. Extensive quantitative and qualitative experiments demonstrate the effectiveness of our method. Ruitong Gan, Chuanchen Luo, Yuxi Wang 0001, Qing Li 0001, Xu-Cheng Yin, Man Zhang 0005, Zhaoxiang Zhang 0001, Junran Peng |
ACM Multimedia | 9 |
| 2024 | Visual Harmony: LLM's Power in Crafting Coherent Indoor Scenes from Images
Genghao Zhang, Yuxi Wang 0001, Chuanchen Luo, Shibiao Xu, Junran Peng, Man Zhang 0005 |
PRCV (6) | 7 |
| 2024 | Bidirectional Knowledge Reconfiguration for Lightweight Point Cloud AnalysisabstractPoint cloud analysis faces computational system overhead, limiting its application on mobile or edge devices. Directly employing small models may result in a significant drop in performance since it is difficult for a small model to adequately capture local structure and global shape information simultaneously, which are essential clues for point cloud analysis. This paper explores feature distillation for lightweight point cloud models. To mitigate the semantic gap between the lightweight student and the cumbersome teacher, we propose bidirectional knowledge reconfiguration (BKR) to distill informative contextual knowledge from the teacher to the student. Specifically, a top-down knowledge reconfiguration and a bottom-up knowledge reconfiguration are developed to inherit diverse local structure information and consistent global shape knowledge from the teacher, respectively. However, due to the farthest point sampling in most point cloud models, the intermediate features between teacher and student are misaligned, deteriorating the feature distillation performance. To eliminate it, we propose a feature mover's distance (FMD) loss based on optimal transportation, which can measure the distance between unordered point cloud features effectively. Extensive experiments conducted on shape classification, part segmentation, and semantic segmentation benchmarks demonstrate the universality and superiority of our method. Peipei Li 0002, Xing Cui, Yibo Hu 0001, Man Zhang 0005, Ting Yao 0003, Tao Mei 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | GaitParsing: Human Semantic Parsing for Gait RecognitionabstractGait recognition is a soft biotechnology to identify pedestrians observed from different camera views based on specific walking patterns. However, various dressing and wearing conditions bring great challenges to realistic gait recognition. Most existing methods take holistic gait silhouette as input and focus on local areas through horizontal strip division or attention map. We consider that this processing may contain mixed or incomplete information about multiple body parts so that gait information is misused or underutilized. In this paper, we propose a parsing-guided framework for gait recognition, namedGaitParsing, which explores human semantic parsing to dissect human body into a set of specific and complete body parts. Correspondingly, a simple yet effective dual-branch feature extraction network is adopted to process holistic gait and distinct body parts. To maximize the use of highly discriminated gait frames, we propose a self-occlusion frame assessment to measure the self-occlusion in a gait sequence. Since there is no human parsing modality in current gait datasets, we further develop a general human parsing pipeline specifically tailored for gait datasets. This single training enables widespread application across various gait datasets. Extensive experiments with ablation analyses demonstrate competitive performance even in the most challenging conditions, e.g., Cloth-Changing (CC+5.9%). Especially, It is gratifying to see that our model can be easily applied to existing methods and significantly outperform the original architecture, even without much modification. Zengbin Wang, Saihui Hou, Man Zhang 0005, Xu Liu 0008, Chunshui Cao, Yongzhen Huang |
IEEE Trans. Multim. | 3 |
| 2023 | DDG-Net: Discriminability-Driven Graph Network for Weakly-supervised Temporal Action LocalizationabstractWeakly-supervised temporal action localization (WTAL) is a practical yet challenging task. Due to large-scale datasets, most existing methods use a network pretrained in other datasets to extract features, which are not suitable enough for WTAL. To address this problem, researchers design several modules for feature enhancement, which improve the performance of the localization module, especially modeling the temporal relationship between snippets. However, all of them omit that ambiguous snippets deliver contradictory information, which would reduce the discriminability of linked snippets. Considering this phenomenon, we propose Discriminability-Driven Graph Network (DDG-Net), which explicitly models ambiguous snippets and discriminative snippets with well-designed connections, preventing the transmission of ambiguous information and enhancing the discriminability of snippet-level representations. Additionally, we propose feature consistency loss to prevent the assimilation of features and drive the graph convolution network to generate more discriminative representations. Extensive experiments on THUMOS14 and ActivityNet1.2 benchmarks demonstrate the effectiveness of DDG-Net, establishing new state-of-the-art results on both datasets. Source code is available at https://github.com/XiaojunTang22/ICCV2023-DDGNet. Junsong Fan, Chuanchen Luo, Zhaoxiang Zhang 0001, Man Zhang 0005, Zongyuan Yang |
ICCV | 5 |
| 2023 | LandmarkGait: Intrinsic Human Parsing for Gait RecognitionabstractGait recognition is an emerging biometric technology for identifying pedestrians based on their unique walking patterns. In past gait recognition, global-based methods are inadequate to meet the growing demand for accuracy, while commonly used part-based methods provided coarse and inaccurate feature representation for specific body parts. Human parsing appears to be a better option for accurately representing specific and complete body parts in gait recognition. However, its practical application in gait recognition is often hindered by missing RGB modality, lack of annotated body parts, and difficulty in balancing parsing quantity and quality. To address this issue, we propose LandmarkGait, an accessible and alternative parsing-based solution for gait recognition. LandmarkGait introduces an unsupervised landmark discovery network to transform the dense silhouette into a finite set of landmarks with remarkable consistency across various conditions. By grouping landmarks subsets corresponding to distinct body part regions, following a reconstruction task and further refinement from high-quality input silhouettes, we can directly obtain fine-grained parsing results from original binary silhouettes in an unsupervised manner. Moreover, we also develop a multi-scale feature extractor that simultaneously captures global and parsing feature representations based on the integrity and flexibility of specific body parts. Extensive experiments demonstrate that our LandmarkGait can extract more stable features and exhibit significant performance improvement under all conditions, especially in various dressing conditions. Code is available at https://github.com/wzb-bupt/LandmarkGait. Zengbin Wang, Saihui Hou, Man Zhang 0005, Xu Liu 0008, Chunshui Cao, Yongzhen Huang, Shibiao Xu |
ACM Multimedia | 3 |
| 2022 | Group Activity Representation Learning with Self-supervised Predictive Coding
Longteng Kong, Zhaofeng He 0001, Man Zhang 0005, Yunzhi Xue |
PRCV (3) | 3 |
| 2019 | Toward practical remote iris recognition: A boosting based framework
Man Zhang 0005, Zhaofeng He 0001, Hui Zhang 0061, Tieniu Tan, Zhenan Sun |
Neurocomputing | 1 |
| 2018 | Adversarial Discriminative Heterogeneous Face RecognitionabstractThe gap between sensing patterns of different face modalities remains a challenging problem in heterogeneous face recognition (HFR). This paper proposes an adversarial discriminative feature learning framework to close the sensing gap via adversarial learning on both raw-pixel space and compact feature space. This framework integrates cross-spectral face hallucination and discriminative feature learning into an end-to-end adversarial network. In the pixel space, we make use of generative adversarial networks to perform cross-spectral face hallucination. An elaborate two-path model is introduced to alleviate the lack of paired images, which gives consideration to both global structures and local textures. In the feature space, an adversarial loss and a high-order variance discrepancy loss are employed to measure the global and local discrepancy between two heterogeneous distributions respectively. These two losses enhance domain-invariant feature learning and modality independent noise removing. Experimental results on three NIR-VIS databases show that our proposed approach outperforms state-of-the-art HFR methods, without requiring of complex network or large-scale training dataset. Lingxiao Song, Man Zhang 0005, Xiang Wu 0001, Ran He 0001 |
AAAI | 2 |
| 2018 | Demographic Analysis from Biometric Data: Achievements, Challenges, and New FrontiersabstractBiometrics is the technique of automatically recognizing individuals based on their biological or behavioral characteristics. Various biometric traits have been introduced and widely investigated, including fingerprint, iris, face, voice, palmprint, gait and so forth. Apart from identity, biometric data may convey various other personal information, covering affect, age, gender, race, accent, handedness, height, weight, etc. Among these, analysis of demographics (age, gender, and race) has received tremendous attention owing to its wide real-world applications, with significant efforts devoted and great progress achieved. This survey first presents biometric demographic analysis from the standpoint of human perception, then provides a comprehensive overview of state-of-the-art advances in automated estimation from both academia and industry. Despite these advances, a number of challenging issues continue to inhibit its full potential. We second discuss these open problems, and finally provide an outlook into the future of this very active field of research by sharing some promising opportunities. Yunlian Sun, Man Zhang 0005, Zhenan Sun, Tieniu Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2017 | Bin-based classifier fusion of iris and face biometrics
Di Miao, Man Zhang 0005, Zhenan Sun, Tieniu Tan, Zhaofeng He 0001 |
Neurocomputing | 2 |
| 2016 | Discriminative Analysis Dictionary LearningabstractDictionary learning (DL) has been successfully applied to various pattern classification tasks in recent years. However, analysis dictionary learning (ADL), as a major branch of DL, has not yet been fully exploited in classification due to its poor discriminability. This paper presents a novel DL method, namely Discriminative Analysis Dictionary Learning (DADL), to improve the classification performance of ADL. First, a code consistent term is integrated into the basic analysis model to improve discriminability. Second, a triplet constraint-based local topology preserving loss function is introduced to capture the discriminative geometrical structures embedded in data. Third, correntropy induced metric is employed as a robust measure to better control outliers for classification. Then, half-quadratic minimization and alternate search strategy are used to speed up the optimization process so that there exist closed-form solutions in each alternating minimization stage. Experiments on several commonly used databases show that our proposed method not only significantly improves the discriminative ability of ADL, but also outperforms state-of-the-art synthesis DL methods. Jun Guo 0008, Yanqing Guo, Xiangwei Kong 0001, Man Zhang 0005, Ran He 0001 |
AAAI | 4 |
| 2016 | Simultaneous Feature and Sample Reduction for Image-Set ClassificationabstractImage-set classification is the assignment of a label to a given image set. In real-life scenarios such as surveillance videos, each image set often contains much redundancy in terms of features and samples. This paper introduces a joint learning method for image-set classification that simultaneously learns compact binary codes and removes redundant samples. The joint objective function of our model mainly includes two parts. The first part seeks a hashing function to generate binary codes that have larger inter-class and smaller intra-class distances. The second one reduces redundant samples with discrete constraints in a low-rank way. A kernel method based on anchor points is further used to reduce sample variations. The proposed discrete objective function is simplified to a series of sub-problems that admit an analytical solution, resulting in a high-quality discrete solution with a low computational cost. Experiments on three commonly used image-set datasets show that the proposed method for the tasks of face recognition from image sets is efficient and effective. Man Zhang 0005, Ran He 0001, Zhenan Sun, Tieniu Tan |
AAAI | 1 |
| 2016 | DeepIris: Learning pairwise filter bank for heterogeneous iris verification
Nianfeng Liu, Man Zhang 0005, Zhenan Sun, Tieniu Tan |
Pattern Recognit. Lett. | 2 |
| 2015 | Iris Texture Description Using Ordinal Co-occurrence Matrix Features
Yasser Chacon-Cabrera, Man Zhang 0005, Eduardo Garea Llano, Zhenan Sun |
CIARP | 2 |
| 2015 | Cross-Modal Subspace Learning via Pairwise ConstraintsabstractIn multimedia applications, the text and image components in a web document form a pairwise constraint that potentially indicates the same semantic concept. This paper studies cross-modal learning via the pairwise constraint and aims to find the common structure hidden in different modalities. We first propose a compound regularization framework to address the pairwise constraint, which can be used as a general platform for developing cross-modal algorithms. For unsupervised learning, we propose a multi-modal subspace clustering method to learn a common structure for different modalities. For supervised learning, to reduce the semantic gap and the outliers in pairwise constraints, we propose a cross-modal matching method based on compound ℓ21 regularization. Extensive experiments demonstrate the benefits of joint text and image modeling with semantically induced pairwise constraints, and they show that the proposed cross-modal methods can further reduce the semantic gap between different modalities and improve the clustering/matching accuracy. Ran He 0001, Man Zhang 0005, Liang Wang 0001, Qiyue Yin |
IEEE Trans. Image Process. | 2 |
| 2014 | The first ICB* competition on iris recognitionabstractIris recognition becomes an important technology in our society. Visual patterns of human iris provide rich texture information for personal identification. However, it is greatly challenging to match intra-class iris images with large variations in unconstrained environments because of noises, illumination variation, heterogeneity and so on. To track current state-of-the-art algorithms in iris recognition, we organized the first ICB* Competition on Iris Recognition in 2013 (or ICIR2013 shortly). In this competition, 8 participants from 6 countries submitted 13 algorithms totally. All the algorithms were trained on a public database (e.g. CASIA-Iris-Thousand [3]) and evaluated on an unpublished database. The testing results in terms of False Non-match Rate (FNMR) when False Match Rate (FMR) is 0.0001 are taken to rank the submitted algorithms. Man Zhang 0005, Jing Liu 0062, Zhenan Sun, Tieniu Tan, Wu Su, Fernando Alonso-Fernandez, Valérian Némesin, Nadia Othman, Koichi Noda, Peihua Li, Edmundo Hoyle, Akanksha Joshi |
IJCB | 1 |
| 2014 | Transform-invariant dictionary learning for face recognitionabstractDictionary learning has important applications in face recognition. However, large transformation variations of face images pose a grand challenge to conventional dictionary learning methods. A large portion of misleading dictionary atoms are usually learned to represent transformation factors, which will cause ambiguity in face recognition. To address this problem, this paper proposes a general framework for transform-invariant basis matrix learning. Specifically, we present a transform-invariant dictionary learning method which explicitly incorporates an appearance consistent error term to the original objective function in dictionary learning. The unified objective function is effectively optimized in an alternating iterative way. An ensemble of aligned images and a discriminative transform-invariant dictionary for sparse coding can be obtained by solving the formulated objective function. Experimental results on two public face databases demonstrate our algorithm's superiority compared with two state-of-the-art dictionary learning methods and the recently proposed transform-invariant PCA method. Shu Zhang 0015, Man Zhang 0005, Ran He 0001, Zhenan Sun |
ICIP | 2 |
| 2011 | Deformable DAISY Matcher for robust iris recognitionabstractIris is rich of texture information for reliable personal identification. However, nonlinear deformation of iris pattern caused by pupil dilation or contraction raises a grand challenge to iris recognition. This paper proposes a novel iris recognition method namely Deformable DAISY Matcher (DDM) for robust iris feature matching. Firstly, dense DAISY descriptors are extracted to represent regional iris features, which are robust against intra-class variations of iris images. Then a set of iris key points are localized on the feature map. Finally deformation tolerant matching strategy is proposed to match corresponding key points of iris images. Experimental results on two iris image databases demonstrate DDM is better than state-of-the-art iris recognition methods. Man Zhang 0005, Zhenan Sun, Tieniu Tan |
ICIP | 1 |