VLDB 2026 Research / reviewers in the wild / expert
Haiyun Guo
dblp:163/0477
· DBLP profile ↗
36ranked-venue papers
7as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 6 first-author · 12 since 2021Artificial intelligence and machine learning · 17 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PASs-MoE: Mitigating Misaligned Co-drift among Router and Experts via Pathway Activation Subspaces for Continual LearningabstractContinual instruction tuning (CIT) requires multimodal large language models (MLLMs) to adapt to a stream of tasks without forgetting prior capabilities.A common strategy is to isolate updates by routing inputs to different LoRA experts.However, existing LoRAbased Mixture-of-Experts (MoE) methods often jointly update the router and experts in an indiscriminate way, causing the router's preferences to co-drift with experts' adaptation pathways and gradually deviate from early-stage input-expert specialization.We term this as Misaligned Co-drift, which blurs expert responsibilities and exacerbates forgetting.To address this, we introduce the pathway activation subspace (PASs), a LoRA-induced subspace that reflects which low-rank pathway directions an input activates in each expert, providing a capability-aligned coordinate system for routing and preservation.Based on PASs, we propose a fixed-capacity PASsbased MoE-LoRA method with two components: PAS-guided Reweighting, which calibrates routing using each expert's pathway activation signals, and PAS-aware Rank Stabilization, which selectively stabilizes rank directions important to previous tasks.Experiments on a CIT benchmark show that our approach consistently outperforms a range of conventional continual learning baselines and MoE-LoRA variants in both accuracy and antiforgetting without adding parameters.Our code is publicly available at https://github.com/ yueluoshuangtian/PASs-MoE. ZhiYan Hou, Haiyun Guo, Haokai Ma, Yandu Sun, Jinqiao Wang |
ACL (1) | 2 |
| 2026 | HB-Mamba: Hierarchical Bi-directional State Space Modeling for LiDAR Semantic Segmentation in Autonomous Drivingabstract3D semantic segmentation remains a pivotal challenge for autonomous driving due to the inherent sparsity of points. Existing CNN-based and Transformer-based methods struggle with either limited receptive fields or quadratic computational complexity. Although some Mamba-based 3D models are designed efficiently with linear complexity, they often overlook the long-term decay problem in Selective State-space Models when processing extremely long sequences in large-scale scenes. In this paper, we propose a Hierarchical Bi-directional Mamba (HB-Mamba) for point cloud semantic segmentation. By decoupling feature extraction into a Global Memory branch and a Local Detail branch, our architecture effectively captures long-range semantics and preserves fine-grained geometric information. Besides, we further introduce a Spatial-Channel Fusion Block to dynamically fuse these multi-scale representations. Experimental results on the nuScenes-Lidarseg benchmark demonstrate that HB-Mamba achieves state-of-the-art performance among Lidar-only methods, reaching 82.8% mIoU on the test set and 81.33% mIoU on the validation set, outperforming the leading transformer-based model PTv3 by 0.1% and 1.01%, respectively. Wei Li 0315, Haiyun Guo, Manli Tao, Honghui Dong, Ming Tang 0001, Jinqiao Wang |
ICMR | 2 |
| 2026 | Continual Instruction Tuning for Large Multimodal ModelsabstractInstruction tuning has become a widely adopted approach for aligning large multimodal models (LMMs) with human intent. It enables multi-task joint training through unified data formats. However, as new vision-language tasks constantly emerge, exhaustive joint training of all tasks becomes impractical. Continual learning offers a more flexible and resource-efficient alternative, enabling incremental training of LMMs on emerging tasks. This study investigates two fundamental questions when applying continual learning to instruction tuning of LMMs: 1) Do LMMs suffer from catastrophic forgetting during continual instruction tuning? 2) Can existing continual learning methods be effectively applied to continual instruction tuning of LMMs? A comprehensive study was conducted to answer these questions. First, we establish the first benchmark for continual instruction tuning of LMMs and reveal the phenomenon of catastrophic forgetting in this setup. Second, we integrate and adapt traditional continual learning approaches to this setting, demonstrating the effectiveness of these strategies to varying degrees in different scenarios. Third, we explore task-similarity dynamics between pairs of vision-language tasks and propose task-similarity-informed regularization and model expansion methods. Experimental results show that our approach can consistently boost the model's performance. Jinghan He, Haiyun Guo, Kuan Zhu, Ming Tang 0001, Jinqiao Wang |
IEEE Trans. Image Process. | 2 |
| 2025 | Cracking the Code of Hallucination in LVLMs with Vision-aware Head DivergenceabstractJinghan He, Kuan Zhu, Haiyun Guo, Junfeng Fang, Zhenglin Hua, Yuheng Jia, Ming Tang, Tat-Seng Chua, Jinqiao Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jinghan He, Kuan Zhu, Haiyun Guo, Junfeng Fang, Zhenglin Hua, Yuheng Jia, Ming Tang 0001, Tat-Seng Chua, Jinqiao Wang |
ACL (1) | 3 |
| 2025 | PhysVLM: Enabling Visual Language Models to Understand Robotic Physical ReachabilityabstractUnderstanding the environment and a robot’s physical reachability is crucial for task execution. While state-of-the-art vision-language models (VLMs) excel in environmental perception, they often generate inaccurate or impractical responses in embodied visual reasoning tasks due to a lack of understanding of robotic physical reachability. To address this issue, we propose a unified representation of physical reachability across diverse robots, i.e., Space-Physical Reachability Map (S-P Map), and PhysVLM, a vision-language model that integrates this reachability information into visual reasoning. Specifically, the S-P Map abstracts a robot’s physical reachability into a generalized spatial representation, independent of specific robot configurations, allowing the model to focus on reachability features rather than robot-specific parameters. Subsequently, PhysVLM extends traditional VLM architectures by incorporating an additional feature encoder to process the S-P Map, enabling the model to reason about physical reachability without compromising its general vision-language capabilities. To train and evaluate PhysVLM, we constructed a large-scale multi-robot dataset, Phys100K, and a challenging benchmark, EQA-phys, which includes tasks for six different robots in both simulated and real-world environments. Experimental results demonstrate that PhysVLM outperforms existing models, achieving a 14% improvement over GPT-4o on EQA-phys and surpassing advanced embodied VLMs such as RoboMamba and SpatialVLM on the RoboVQA-val and OpenEQA benchmarks. Additionally, the S-P Map shows strong compatibility with various VLMs, and its integration into GPT-4o-mini yields a 7.1% performance improvement. Manli Tao, Chaoyang Zhao, Haiyun Guo, Honghui Dong, Ming Tang 0001, Jinqiao Wang |
CVPR | 4 |
| 2025 | FOCUS: Fine-grained Optimization with Semantic Guided Understanding for Pedestrian Attributes RecognitionabstractPedestrian attribute recognition (PAR) is a fundamental perception task in intelligent transportation and security. To tackle this fine-grained task, most existing methods focus on extracting regional features to enrich attribute information. However, a regional feature is typically used to predict a fixed set of pre-defined attributes in these methods, which limits the performance and practicality in two aspects: 1) Regional features may compromise fine-grained patterns unique to certain attributes in favor of capturing common characteristics shared across attributes. 2) Regional features cannot generalize to predict unseen attributes in the test time. In this paper, we propose the Fine-grained Optimization with semantiC gUided underStanding (FOCUS) approach for PAR, which adaptively extracts fine-grained attribute-level features for each attribute individually, regardless of whether the attributes are seen or not during training. Specifically, we propose the Multi-Granularity Mix Tokens (MGMT) to capture latent features at varying levels of visual granularity, thereby enriching the diversity of the extracted information. Next, we introduce the Attribute-guided Visual Feature Extraction (AVFE) module, which leverages textual attributes as queries to retrieve their corresponding visual attribute features from the Mix Tokens using a cross-attention mechanism. To ensure that textual attributes focus on the appropriate Mix Tokens, we further incorporate a Region-Aware Contrastive Learning (RACL) method, encouraging attributes within the same region to share consistent attention maps. Extensive experiments on PA100K, PETA, and RAPv1 datasets demonstrate the effectiveness and strong generalization ability of our method. Hongyan An, Kuan Zhu, Haiyun Guo, Chaoyang Zhao, Ming Tang 0001, Jinqiao Wang |
ICME | 4 |
| 2025 | Semantic-aware Fine-grained Point Augmentation for 3D Multi-modal Object Detectionabstract3D object detection aims to locate and recognize the object from the point cloud, which is a meaningful and foundation task in autonomous driving. However, the sparsity of the point cloud poses a significant challenge for this task, especially for distant and small objects. Existing methods employ depth estimation networks to generate pseudo points for improving the point density, but this introduces significant computational costs and noise, limiting performance gains. In this paper, we propose the Semantic-aware Fine-grained Point Augmentation (SFPA) approach for 3D object detection, which simultaneously enriches high-quality point clouds and filters noisy points, and incorporates multi-modal feature fusion to enhance detection performance. Specifically, we utilize a semantic segmentation model to generate object masks from RGB images and refine dense depth estimation maps, derived from sparse LiDAR points and RGB images, using these foreground object masks. Subsequently, high-quality pseudo point clouds, concentrated solely on foreground objects, are generated by projecting the refined dense depth maps back to 3D coordinates. Furthermore, we also employ the projection matrix as an alignment strategy to concatenate or add dense RGB features with point features, further improving detection performance for extremely sparse objects. Experimental results demonstrate that our method achieves state-of-the-art performance on KITTI 3D object detection leaderboard, i.e., 95.44%, 88.18%, 85.53% for the Car category at the easy, medium, and hard levels, respectively. Wei Li 0315, Kuan Zhu, Haiyun Guo, Honghui Dong, Jinqiao Wang |
ICME | 3 |
| 2025 | Referring Expression Instance Retrieval and A Strong End-to-End BaselineabstractText-Image Retrieval (TIR) retrieves a target image from a gallery based on an image-level description, while Referring Expression Comprehension (REC) localizes a target object within a given image using an instance-level description. However, real-world applications often present more complex demands. Users typically query an instance-level description across a large gallery and expect to receive both relevant image and the corresponding instance location. In such scenarios, TIR struggles with fine-grained descriptions and object-level localization, while REC is limited in its ability to efficiently search large galleries and lacks an effective ranking mechanism. In this paper, we introduce a new task called Referring Expression Instance Retrieval (REIR), which supports both instance-level retrieval and localization based on fine-grained referring expressions. First, we propose a large-scale benchmark for REIR, named REIRCOCO, constructed by prompting advanced vision-language models to generate high quality referring expressions for instances in the MSCOCO and RefCOCO datasets. Second, we present a baseline method, Contrastive Language Instance Alignment with Relation Experts (CLARE), which employs a dual-stream architecture to address REIR in an end-to-end manner. Given a referring expression, the textual branch encodes it into a query embedding, enhanced by a Mix of Relation Experts (MORE) module designed to better capture inter-instance relationships. The visual branch detects candidate objects and extracts their instance-level visual features. The most similar candidate to the query is selected for bounding box prediction. CLARE is first trained on object detection and REC datasets to establish initial grounding capabilities, then optimized via Contrastive Language Instance Alignment (CLIA) for improved retrieval across images. Experimental results demonstrate that CLARE outperforms existing methods on the REIR benchmark and generalizes well to both TIR and REC tasks, showcasing its effectiveness and versatility. Xiangzhao Hao, Kuan Zhu, Haiyun Guo, Ming Tang 0001, Jinqiao Wang |
ACM Multimedia | 4 |
| 2025 | MaSA: Mamba-Based Global Feature Selective Aggregator for Efficient Lane Detection
La Zhang, Haiyun Guo, Chaoyang Zhao, Jinqiao Wang |
PRCV (3) | 3 |
| 2024 | WaveMo: Learning Wavefront Modulations to See Through ScatteringabstractImaging through scattering media is a fundamental and pervasive challenge infields ranging from medical diagnos-tics to astronomy. A promising strategy to overcome this challenge is wavefront modulation, which induces measure-ment diversity during image acquisition. Despite its importance, designing optimal wavefront modulations to image through scattering remains under-explored. This paper in-troduces a novel learning-based framework to address the gap. Our approach jointly optimizes wavefront modulations and a computationally lightweight feedforward “proxy” re-construction network. This network is trained to recover scenes obscured by scattering, using measurements that are modified by these modulations. The learned modulations produced by our framework generalize effectively to un-seen scattering scenarios and exhibit remarkable versatility. During deployment, the learned modulations can be decou-pled from the proxy network to augment other more computationally expensive restoration algorithms. Through ex-tensive experiments, we demonstrate our approach signifi-cantly advances the state of the art in imaging through scat-tering media. Our project webpage is at https://wavemo-2024.github.io/. Mingyang Xie, Haiyun Guo, Brandon Yushan Feng, Lingbo Jin, Ashok Veeraraghavan, Christopher A. Metzler |
CVPR | 2 |
| 2024 | SEEKR: Selective Attention-Guided Knowledge Retention for Continual Learning of Large Language ModelsabstractContinual learning (CL) is crucial for language models to dynamically adapt to the evolving real-world demands.To mitigate the catastrophic forgetting problem in CL, data replay has been proven a simple and effective strategy, and the subsequent data-replay-based distillation can further enhance the performance.However, existing methods fail to fully exploit the knowledge embedded in models from previous tasks, resulting in the need for a relatively large number of replay samples to achieve good results.In this work, we first explore and emphasize the importance of attention weights in knowledge retention, and then propose a SElective attEntion-guided Knowledge Retention method (SEEKR) for data-efficient replay-based continual learning of large language models (LLMs).Specifically, SEEKR performs attention distillation on the selected attention heads for finer-grained knowledge retention, where the proposed forgettabilitybased and task-sensitivity-based measures are used to identify the most valuable attention heads.Experimental results on two continual learning benchmarks for LLMs demonstrate the superiority of SEEKR over the existing methods on both performance and efficiency.Explicitly, SEEKR achieves comparable or even better performance with only 1/10 of the replayed data used by other methods, and reduces the proportion of replayed data to 1%.The code is available at https: //github.com/jinghan1he/SEEKR. Jinghan He, Haiyun Guo, Kuan Zhu, Ming Tang 0001, Jinqiao Wang |
EMNLP | 2 |
| 2024 | AAformer: Auto-Aligned Transformer for Person Re-IdentificationabstractIn person re-identification (re-ID), extracting part-level features from person images has been verified to be crucial to offer fine-grained information. Most of the existing CNN-based methods only locate the human parts coarsely, or rely on pretrained human parsing models and fail in locating the identifiable nonhuman parts (e.g., knapsack). In this article, we introduce an alignment scheme in transformer architecture for the first time and propose the auto-aligned transformer (AAformer) to automatically locate both the human parts and nonhuman ones at patch level. We introduce the "Part tokens ([PART]s)," which are learnable vectors, to extract part features in the transformer. A [PART] only interacts with a local subset of patches in self-attention and learns to be the part representation. To adaptively group the image patches into different subsets, we design the auto-alignment. Auto-alignment employs a fast variant of optimal transport (OT) algorithm to online cluster the patch embeddings into several groups with the [PART]s as their prototypes. AAformer integrates the part alignment into the self-attention and the output [PART]s can be directly used as part features for retrieval. Extensive experiments validate the effectiveness of [PART]s and the superiority of AAformer over various state-of-the-art methods. Kuan Zhu, Haiyun Guo, Shiliang Zhang, Yaowei Wang 0001, Jing Liu 0001, Jinqiao Wang, Ming Tang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Instance-Proxy Loss for Semi-supervised Learning with Coarse Labels
Qinghai Miao, Haiyun Guo, Min Huang 0009, Jinqiao Wang |
PRCV (12) | 3 |
| 2023 | Bi-Level Implicit Semantic Data Augmentation for Vehicle Re-IdentificationabstractVehicle re-identification (Re-ID) aims at finding the target vehicle identity from multi-camera surveillance videos, which plays an important role in the intelligent transportation system (ITS). It suffers from the subtle discrepancy among vehicles from the same vehicle model and large variation across different viewpoints of the same vehicle. To enhance the robustness of Re-ID models, many methods exploit additional detection or segmentation models to extract discriminative local features. Some others employ data-driven methods to enrich the diversity of the training data, such as the data augmentation and 3D-based data generation, so that the Re-ID model can obtain stronger robustness against intra-class variations. However, these methods either rely on extra annotations or greatly increase the computational cost. In this paper, we propose the Bi-level Implicit semantic Data Augmentation (BIDA) framework to solve this problem from two aspects. (1) We implicitly augment the images semantically in the feature space according to the identity-level and superclass-level intra-class variations, which can generate more diverse semantic augmentations beyond the intra-identity variations. (2) We introduce the similarity ranking constraints on the augmented training set by extending the sample-wise triplet loss to the distribution-wise one, which can effectively reduce meaningless semantic transformations and improve the discrimination of the feature. We conduct extensive experiments on VeRi-776, VehicleID and Cityflow benchmarks to reveal the effectiveness of our method. And we achieve new state-of-the-art performance on VeRi-776. Wei Li 0315, Haiyun Guo, Honghui Dong, Ming Tang 0001, Jinqiao Wang |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Pseudo Label Rectification With Joint Camera Shift Adaptation and Outlier Progressive Recycling for Unsupervised Person Re-IdentificationabstractPerson re-identification (re-ID) has many applications in intelligent transportation systems. Clustering-based methods, which alternate between the generation of pseudo labels via clustering and the optimization of the feature extractor, have obtained leading performance in unsupervised person re-ID. But there are still two issues not well addressed: 1) Most methods measure the feature similarity without considering the domain shift between cameras, degrading the clustering performance. 2) Outliers, which usually correspond to hard samples with large discrepancy from other images of the identical person, are in most cases directly excluded from the network training. To tackle the above issues, this paper proposes a plug-and-play pseudo label rectification framework, which jointly utilizes CAmera Shift adapTation module and Outlier progressive Recycling strategy ($CASTOR$) to improve the quality of pseudo labels from both pre-clustering and post-clustering. Specifically, we first compute the camera similarity of two samples by utilizing a pretrained camera classification network and subtract the feature similarity by the camera similarity, the value of which is weighted in an exponential decay manner throughout the network training, in order to adaptively remedy the adverse impact of inter-camera distribution shift upon clustering. Besides, we carefully design an outlier progressive recycling strategy to reassign part of the outliers into the clustered groups to make full use of the useful information of outliers. Extensive experiments on three large scale unsupervised and unsupervised domain adaptive (UDA) person re-ID benchmarks validate the effectiveness of$CASTOR$and its wide compatibility with the state-of-the-art clustering-based methods. Mingyuan Xu, Haiyun Guo, Yuheng Jia, Zhitao Dai, Jinqiao Wang |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Learning Semantics-Consistent Stripes With Self-Refinement for Person Re-IdentificationabstractAligning human parts automatically is one of the most challenging problems for person re-identification (re-ID). Recently, the stripe-based methods, which equally partition the person images into the fixed stripes for aligned representation learning, have achieved great success. However, the stripes with fixed height and position cannot well handle the misalignment problems caused by inaccurate detection and occlusion and may introduce much background noise. In this article, we aim at learning adaptive stripes with foreground refinement to achieve pixel-level part alignment by only using person identity labels for person re-ID and make two contributions. 1) A semantics-consistent stripe learning method (SCS). Given an image, SCS partitions it into adaptive horizontal stripes and each stripe is corresponding to a specific semantic part. Specifically, SCS iterates between two processes: i) clustering the rows to human parts or background to generate the pseudo-part labels of rows and ii) learning a row classifier to partition a person image, which is supervised by the latest pseudo-labels. This iterative scheme guarantees the accuracy of the learned image partition. 2) A self-refinement method (SCS+) to remove the background noise in stripes. We employ the above row classifier to generate the probabilities of pixels belonging to human parts (foreground) or background, which is called the class activation map (CAM). Only the most confident areas from the CAM are assigned with foreground/background labels to guide the human part refinement. Finally, by intersecting the semantics-consistent stripes with the foreground areas, SCS+ locates the human parts at pixel-level, obtaining a more robust part-aligned representation. Extensive experiments validate that SCS+ sets the new state-of-the-art performance on three widely used datasets including Market-1501, DukeMTMC-reID, and CUHK03-NP. Kuan Zhu, Haiyun Guo, Songyan Liu, Jinqiao Wang, Ming Tang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | PASS: Part-Aware Self-Supervised Pre-Training for Person Re-Identification
Kuan Zhu, Haiyun Guo, Tianyi Yan, Yousong Zhu, Jinqiao Wang, Ming Tang 0001 |
ECCV (14) | 2 |
| 2022 | Graph Neural Networks Based Multi-granularity Feature Representation Learning for Fine-Grained Visual Categorization
Haiyun Guo, Qinghai Miao, Min Huang 0009, Jinqiao Wang |
MMM (2) | 2 |
| 2022 | Multi-Granularity Mutual Learning Network for Object Re-IdentificationabstractObject re-identification (re-ID), which is key and fundamental technology for intelligent transportation systems, is a challenging task including person re-ID and vehicle re-ID. It aims to retrieve a given target object from the gallery images captured by different cameras. In this task, it is necessary to extract fine-grained and discriminative features to deal with complex inter-class and intra-class variations caused by the changes of camera viewpoints and object poses. Existing methods focus on learning discriminative local features to improve the re-ID performance. Some state-of-the-art methods use key point detection model to locate local features, which also increases the additional computational cost as side effect. Another type of method focuses on how to learn features of different granularity from rigid stripes of different scales. However, there is little attention paid to how to effectively coalesce multi-granularity features without additional calculation cost. To tackle this issue, this paper proposes the Multi-granularity Mutual Learning Network (MMNet) and makes two contributions. 1) We introduce the multi-granularity jigsaw puzzle module into object re-ID to impel the network to learn local discriminative features from multiple visual granularities by breaking spatial correlation in original images. 2) We propose a parameter-free multi-scale feature reconstruction module to facilitate mutual learning of features at multiple grain levels, thereby both global features and local features have strong representation capabilities. Extensive experiments demonstrate the effectiveness of our proposed modules and the superiority of our method over various state-of-the-art methods on both person and vehicle re-ID benchmarks. Mingfei Tu, Kuan Zhu, Haiyun Guo, Qinghai Miao, Chaoyang Zhao, Guibo Zhu, Honglin Qiao, Gaopan Huang, Ming Tang 0001, Jinqiao Wang |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Hybrid Modality Metric Learning for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (Re-ID) has received increasing research attention for its great practical value in night-time surveillance scenarios. Due to the large variations in person pose, viewpoint, and occlusion in the same modality, as well as the domain gap brought by heterogeneous modality, this hybrid modality person matching task is quite challenging. Different from the metric learning methods for visible person re-ID, which only pose similarity constraints on class level, an efficient metric learning approach for visible-infrared person Re-ID should take both the class-level and modality-level similarity constraints into full consideration to learn sufficiently discriminative and robust features. In this article, the hybrid modality is divided into two types, within modality and cross modality. We first fully explore the variations that hinder the ranking results of visible-infrared person re-ID and roughly summarize them into three types: within-modality variation, cross-modality modality-related variation, and cross-modality modality-unrelated variation. Then, we propose a comprehensive metric learning framework based on four kinds of paired-based similarity constraints to address all the variations within and cross modality. This framework focuses on both class-level and modality-level similarity relationships between person images. Furthermore, we demonstrate the compatibility of our framework with any paired-based loss functions by giving detailed implementation of combing it with triplet loss and contrastive loss separately. Finally, extensive experiments of our approach on SYSU-MM01 and RegDB demonstrate the effectiveness and superiority of our proposed metric learning framework for visible-infrared person Re-ID. La Zhang, Haiyun Guo, Kuan Zhu, Honglin Qiao, Gaopan Huang, Huichen Zhang, Jian Sun 0003, Jinqiao Wang |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2021 | Unsupervised cycle-consistent person pose transfer
Songyan Liu, Haiyun Guo, Kuan Zhu, Jinqiao Wang, Ming Tang 0001 |
Neurocomputing | 2 |
| 2020 | Adaptive Variance Based Label Distribution Learning for Facial Age Estimation
Xin Wen 0005, Biying Li, Haiyun Guo, Zhiwei Liu 0004, Guosheng Hu, Ming Tang 0001, Jinqiao Wang |
ECCV (23) | 3 |
| 2020 | Identity-Guided Human Semantic Parsing for Person Re-identification
Kuan Zhu, Haiyun Guo, Zhiwei Liu 0004, Ming Tang 0001, Jinqiao Wang |
ECCV (3) | 2 |
| 2020 | A novel data augmentation scheme for pedestrian detection with attribute preserving GAN
Songyan Liu, Haiyun Guo, Jian-Guo Hu, Xu Zhao 0003, Chaoyang Zhao, Tong Wang 0015, Yousong Zhu, Jinqiao Wang, Ming Tang 0001 |
Neurocomputing | 2 |
| 2019 | Semantic Alignment: Finding Semantically Consistent Ground-Truth for Facial Landmark DetectionabstractRecently, deep learning based facial landmark detection has achieved great success. Despite this, we notice that the semantic ambiguity greatly degrades the detection performance. Specifically, the semantic ambiguity means that some landmarks (e.g. those evenly distributed along the face contour) do not have clear and accurate definition, causing inconsistent annotations by annotators. Accordingly, these inconsistent annotations, which are usually provided by public databases, commonly work as the ground-truth to supervise network training, leading to the degraded accuracy. To our knowledge, little research has investigated this problem. In this paper, we propose a novel probabilistic model which introduces a latent variable, i.e. the `real' ground-truth which is semantically consistent, to optimize. This framework couples two parts (1) training landmark detection CNN and (2) searching the `real' ground-truth. These two parts are alternatively optimized: the searched `real' ground-truth supervises the CNN training; and the trained CNN assists the searching of `real' ground-truth. In addition, to recover the unconfidently predicted landmarks due to occlusion and low quality, we propose a global heatmap correction unit (GHCU) to correct outliers by considering the global face shape as a constraint. Extensive experiments on both image-based (300W and AFLW) and video-based (300-VW) databases demonstrate that our method effectively improves the landmark detection accuracy and achieves the state of the art performance. Zhiwei Liu 0004, Xiangyu Zhu 0001, Guosheng Hu, Haiyun Guo, Ming Tang 0001, Zhen Lei 0001, Neil Robertson 0002, Jinqiao Wang |
CVPR | 4 |
| 2019 | Cascade Attention Network for Person Re-IdentificationabstractPerson re-identification is a challenging task due to the viewpoint, illumination and pose variations. Recent works focus on extracting part-level features to offer beneficial fine-grained information. However, the part misalignment as well as the multi-stage training process limits their performance. Inspired by the human visual attention mechanism, this paper builds a cascade attention network(CAN) to learn the discriminative person features in a coarse-to-fine manner. Firstly, we employ the human semantic parsing module to generate coarse-grained part-level attention, which corresponds to the division of human body parts and can effectively filter the background noise. Then, to extract the local detailed features within each part, we introduce spatial-channel attention module to generate fine-grained pixel-level attention, which can further highlight the distinctive characteristics and repress the irrelevant ones. Finally, we can obtain an efficient person feature descriptor by combining both the global and local features. The whole learning process is conducted end-to-end. Experimental results show that the proposed method not only considerably outperforms its counter part but also achieves competitive performance on Market-1501 and DukeMTMC. Haiyun Guo, Huiyao Wu, Chaoyang Zhao, Huichen Zhang, Jinqiao Wang, Hanqing Lu |
ICIP | 1 |
| 2019 | Elite Loss for scene text detection
Xu Zhao 0003, Chaoyang Zhao, Haiyun Guo, Yousong Zhu, Ming Tang 0001, Jinqiao Wang |
Neurocomputing | 3 |
| 2019 | Two-Level Attention Network With Multi-Grain Ranking Loss for Vehicle Re-IdentificationabstractVehicle re-identification (re-ID) aims to identify the same vehicle across multiple non-overlapping cameras, which is rather a challenging task. On the one hand, subtle changes in viewpoint and illumination condition can make the same vehicle look much different. On the other hand, different vehicles, even different vehicle models, may look quite similar. In this paper, we propose a novel Two-level Attention network supervised by a Multi-grain Ranking loss (TAMR) to learn an efficient feature embedding for the vehicle re-ID task. The two-level attention network consisting of hard part-level attention and soft pixel-level attention can adaptively extract discriminative features from the visual appearance of vehicles. The former one is designed to localize the salient vehicle parts, such as windscreen and car head. The latter one gives an additional attention refinement at pixel level to focus on the distinctive characteristics within each part. In addition, we present a multi-grain ranking loss to further enhance the discriminative ability of learned features. We creatively take the multi-grain relationship between vehicles into consideration. Thus, not only the discrimination between different vehicles but also the distinction between different vehicle models is constrained. Finally, the proposed network can learn a feature space, where both intra-class compactness and inter-class discrimination are well guaranteed. Extensive experiments demonstrate the effectiveness of our approach and we achieve state-of-the-art results on two challenging datasets, including VehicleID and Vehicle-1M. Haiyun Guo, Kuan Zhu, Ming Tang 0001, Jinqiao Wang |
IEEE Trans. Image Process. | 1 |
| 2019 | Attention CoupleNet: Fully Convolutional Attention Coupling Network for Object DetectionabstractThe field of object detection has made great progress in recent years. Most of these improvements are derived from using a more sophisticated convolutional neural network. However, in the case of humans, the attention mechanism, global structure information, and local details of objects all play an important role for detecting an object. In this paper, we propose a novel fully convolutional network, named as Attention CoupleNet, to incorporate the attention-related information and global and local information of objects to improve the detection performance. Specifically, we first design a cascade attention structure to perceive the global scene of the image and generate class-agnostic attention maps. Then the attention maps are encoded into the network to acquire object-aware features. Next, we propose a unique fully convolutional coupling structure to couple global structure and local parts of the object to further formulate a discriminative feature representation. To fully explore the global and local properties, we also design different coupling strategies and normalization ways to make full use of the complementary advantages between the global and local information. Extensive experiments demonstrate the effectiveness of our approach. We achieve state-of-the-art results on all three challenging data sets, i.e., a mAP of 85.7% on VOC07, 84.3% on VOC12, and 35.4% on COCO. Codes are publicly available at https://github.com/tshizys/CoupleNet. Yousong Zhu, Chaoyang Zhao, Haiyun Guo, Jinqiao Wang, Xu Zhao 0003, Hanqing Lu |
IEEE Trans. Image Process. | 3 |
| 2018 | Learning Coarse-to-Fine Structured Feature Embedding for Vehicle Re-IdentificationabstractVehicle re-identification (re-ID) is to identify the same vehicle across different cameras. It’s a significant but challenging topic, which has received little attention due to the complex intra-class and inter-class variation of vehicle images and the lack of large-scale vehicle re-ID dataset. Previous methods focus on pulling images from different vehicles apart but neglect the discrimination between vehicles from different vehicle models, which is actually quite important to obtain a correct ranking order for vehicle re-ID. In this paper, we learn a structured feature embedding for vehicle re-ID with a novel coarse-to-fine ranking loss to pull images of the same vehicle as close as possible and achieve discrimination between images from different vehicles as well as vehicles from different vehicle models. In the learnt feature space, both intra-class compactness and inter-class distinction are well guaranteed and the Euclidean distance between features directly reflects the semantic similarity of vehicle images. Furthermore, we build so far the largest vehicle re-ID dataset "Vehicle-1M," which involves nearly 1 million images captured in various surveillance scenarios. Experimental results on "Vehicle-1M" and "VehicleID" demonstrate the superiority of our proposed approach. Haiyun Guo, Chaoyang Zhao, Zhiwei Liu 0004, Jinqiao Wang, Hanqing Lu |
AAAI | 1 |
| 2017 | Deep embedding network for robust age estimationabstractEstimating age through a single facial image is a classic and challenging topic in computer vision. Since facial images of the same age vary considerably, while those from different ages may look very similar. To address these problems, we propose an end-to-end deep embedding neural network for robust age estimation. Specifically, we jointly use classification loss and triplet-based ranking loss to train a deep embedding network, which maps the input facial images into an embedding metric space where features of the same age are compact and those from different ages are pushed away. Thus the deep embedding network can learn more discriminative features and improves the performance for age estimation. Additionally, to accelerate the convergence of the network, we adopt an online hard negative mining strategy during the triplet loss computation. Experimental results on public datasets MORPH II and FG-NET show the superiority of our approach compared to the state-of-the-art. Yating He, Min Huang 0009, Qinghai Miao, Haiyun Guo, Jinqiao Wang |
ICIP | 4 |
| 2016 | Scale-Adaptive Deconvolutional Regression Network for Pedestrian Detection
Yousong Zhu, Jinqiao Wang, Chaoyang Zhao, Haiyun Guo, Hanqing Lu |
ACCV (2) | 4 |
| 2016 | Multiple deep features learning for object retrieval in surveillance videosabstractEfficient indexing and retrieving objects of interest from large‐scale surveillance videos are a significant and challenging topic. In this study, the authors present an effective multiple deep features learning approach for object retrieval in surveillance videos. Based on the discriminative convolutional neural network (CNN), they can learn multiple deep features to comprehensively describe the visual object. To be specific, they utilise the CNN model pre‐trained on ImageNet ILSVRC12 and fine‐tuned on our dataset to abstract structure information. In addition, they train another CNN model supervised by 11 colour names to deliver the colour information. To improve the retrieval performance, the deep features are encoded into short binary codes by locality‐sensitive hash and fused to fast retrieve the object of interest. Retrieval experiments are performed on a dataset of 100k objects extracted from multi‐camera surveillance videos. Comparison results with other common visual features show the effectiveness of the proposed approach. Haiyun Guo, Jinqiao Wang, Hanqing Lu |
IET Comput. Vis. | 1 |
| 2016 | Multi-View 3D Object Retrieval With Deep Embedding NetworkabstractIn multi-view 3D object retrieval, each object is characterized by a group of 2D images captured from different views. Rather than using hand-crafted features, in this paper, we take advantage of the strong discriminative power of convolutional neural network to learn an effective 3D object representation tailored for this retrieval task. Specifically, we propose a deep embedding network jointly supervised by classification loss and triplet loss to map the high-dimensional image space into a low-dimensional feature space, where the Euclidean distance of features directly corresponds to the semantic similarity of images. By effectively reducing the intra-class variations while increasing the inter-class ones of the input images, the network guarantees that similar images are closer than dissimilar ones in the learned feature space. Besides, we investigate the effectiveness of deep features extracted from different layers of the embedding network extensively and find that an efficient 3D object representation should be a tradeoff between global semantic information and discriminative local characteristics. Then, with the set of deep features extracted from different views, we can generate a comprehensive description for each 3D object and formulate the multi-view 3D object retrieval as a set-to-set matching problem. Extensive experiments on SHREC'15 data set demonstrate the superiority of our proposed method over the previous state-of-the-art approaches with over 12% performance improvement. Haiyun Guo, Jinqiao Wang, Yue Gao 0002, Jianqiang Li 0002, Hanqing Lu |
IEEE Trans. Image Process. | 1 |
| 2015 | Learning deep compact descriptor with bagging auto-encoders for object retrievalabstractContent based object retrieval across large scale surveillance video dataset is a significant and challenging task, in which learning an effective compact object descriptor plays a critical role. In this paper, we propose an efficient deep compact descriptor with bagging auto-encoders. Specifically, we take advantage of discriminative CNN to extract efficient deep features, which not only involve rich semantic information but also can filter background noise. Besides, to boost the retrieval speed, auto-encoders are used to map the high-dimensional real-valued CNN features into short binary codes. Considering the instability of auto-encoder, we adopt a bagging strategy to fuse multiple auto-encoders to reduce the generalization error, thus further improving the retrieval accuracy. In addition, bagging is easy for parallel computing, so retrieval efficiency can be guaranteed. Retrieval experimental results on the dataset of 100k visual objects extracted from multi-camera surveillance videos demonstrate the effectiveness of the proposed deep compact descriptor. Haiyun Guo, Jinqiao Wang, Hanqing Lu |
ICIP | 1 |
| 2015 | Learning Multi-view Deep Features for Small Object Retrieval in Surveillance ScenariosabstractWith the explosive growth of surveillance videos, object retrieval has become a significant task for security monitoring. However, visual objects in surveillance videos are usually of small size with complex light conditions, view changes and partial occlusions, which increases the difficulty level of efficiently retrieving objects of interest in a large-scale dataset. Although deep features have achieved promising results on object classification and retrieval and have been verified to contain rich semantic structure property, they lack of adequate color information, which is as crucial as structure information for effective object representation. In this paper, we propose to leverage discriminative Convolutional Neural Network (CNN) to learn deep structure and color feature to form an efficient multi-view object representation. Specifically, we utilize CNN trained on ImageNet to abstract rich semantic structure information. Meanwhile, we propose a CNN model supervised by 11 color names to extract deep color features. Compared with traditional color descriptors, deep color features can capture the common color property across different illumination conditions. Then, the complementary multi-view deep features are encoded into short binary codes by Locality-Sensitive Hash (LSH) and fused to retrieve objects. Retrieval experiments are performed on a dataset of 100k objects extracted from multi-camera surveillance videos. Comparison results with several popular visual descriptors show the effectiveness of the proposed approach. Haiyun Guo, Jinqiao Wang, Min Xu 0001, Zhengjun Zha, Hanqing Lu |
ACM Multimedia | 1 |