VLDB 2026 Research / reviewers in the wild / expert
Tao Wang 0011
dblp:12/5838-11
· DBLP profile ↗
71ranked-venue papers
13as first author
49since 2021 · last 2026
0000-0003-2369-2129ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 12 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 31 · 7 first-author · 20 since 2021Databases, data management, data science and information retrieval · 8 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Computer networks · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorSecurity and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DrawMotion: Generating 3D Human Motions by Freehand Drawing
Tao Wang 0011, Lei Jin 0003, Qiaozhi He, Jiaming Chu, Yu Cheng 0009, Junliang Xing, Jian Zhao 0006, Shuicheng Yan, Li Wang 0039 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | GrassNet: State space model meets graph neural network
Gongpei Zhao, Tao Wang 0011, Yi Jin 0001, Congyan Lang, Yidong Li, Haibin Ling |
Pattern Recognit. | 2 |
| 2026 | SynSP++: General Pose Sequences Refinement via Synergy of Smoothness and Precision
Lei Jin 0003, Tao Wang 0011, Junliang Xing, Jian Zhao 0006, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | StickMotion: Generating 3D Human Motions by Drawing a StickmanabstractText-to-motion generation, which translates textual descriptions into human motions, has been challenging in accurately capturing detailed user-imagined motions from simple text inputs. This paper introduces StickMotion, an efficient diffusion-based network designed for multi-condition scenarios, which generates desired motions based on traditional text and our proposed stickman conditions for global and local control of these motions, respectively. We address the challenges introduced by the user-friendly stickman from three perspectives: 1) Data generation. We develop an algorithm to generate hand-drawn stickmen automatically across different dataset formats. 2) Multi-condition fusion. We propose a multi-condition module that integrates into the diffusion process and obtains outputs of all possible condition combinations, reducing computational complexity and enhancing StickMotion’s performance compared to conventional approaches with the self-attention module. 3) Dynamic supervision. We empower StickMotion to make minor adjustments to the stickman’s position within the output sequences, generating more natural movements through our proposed dynamic supervision strategy. Through quantitative experiments and user studies, sketching stickmen saves users about 51.5% of their time generating motions consistent with their imagination. Our codes, demos, and relevant data will be released in https:// github.com/InvertedForest/StickMotion. Tao Wang 0011, Qiaozhi He, Jiaming Chu, Ling Qian, Yu Cheng 0009, Junliang Xing, Jian Zhao 0006, Lei Jin 0003 |
CVPR | 1 |
| 2025 | A Hubness Perspective on Representation Learning for Graph-Based Multi-View ClusteringabstractRecent graph-based multi-view clustering (GMVC) methods typically encode view features into high-dimensional spaces and construct graphs based on distance similarity. However, the high dimensionality of the embeddings often leads to the hubness problem, where a few points repeatedly appear in the nearest neighbor lists of other points. We show that this negatively impacts the extracted graph structures and message passing, thus degrading clustering performance. To the best of our knowledge, we are the first to highlight the detrimental effect of hubness in GMVC methods and introduce the hubREP (hub-aware Representation Embedding and Pairing) framework. Specifically, we propose a simple yet effective encoder that reduces hubness while preserving neighborhood topology within each view. Additionally, we propose a hub-aware pairing module to maintain structure consistency across views, efficiently enhancing the view-specific representations. The proposed hubREP is lightweight compared to the conventional autoencoders used in state-of-the-art GMVC methods and can be integrated into existing GMVC methods that mostly focus on novel fusion mechanisms, further boosting their performance. Comprehensive experiments performed on eight benchmarks confirm the superiority of our method. The code is available at https://github.com/zmxu196/hubREP. Zheming Xu, Congyan Lang, Tao Wang 0011, Yidong Li, Michael Kampffmeyer |
CVPR | 4 |
| 2025 | CFF: Coarse-to-Fine-to-Fusion Semantic Prototype Generation for Zero-Shot ClassificationabstractZero-Shot Learning focuses on recognizing images from unseen classes with the model trained only on seen classes and auxiliary information. Auxiliary information represents the semantic concepts of classes and is crucial for unseen class generalization. Existing works have tried human-annotated attributes, word embedding, or texts as auxiliary information. However, these auxiliary information is semantically insufficient and visually misaligned, constraining model performance. In this work, we propose the Coarse-to-Fine-to-Fusion prototype generation network (CFF). To obtain vision-oriented text corpora, we design the Coarse-to-Fine Text Generation (CFTG) paradigm, utilizing large language models to generate coarse- and fine-grained texts. For text embedding and fusion, we propose the Semantic Prototype Generation (SPG) module, fusing learnable prompts with coarse- and fine-grained embedding, enabling fine-grained and fusion prototype generation. Moreover, we propose the Visual-Semantic Alignment (VSA) loss to narrow the domain gap. Extensive experiments on AWA2, SUN, and CUB datasets demonstrate the effectiveness of our method. Xuanwen Su, Tengfei Liang, Yi Jin 0001, Tao Wang 0011, Yidong Li |
ICME | 5 |
| 2025 | Tree of Prompts: Aligning Hierarchical Visual Prior for Continual Generalized Category DiscoveryabstractContinual Generalized Category Discovery (C-GCD) aims to incrementally identify both known and novel classes from unlabeled data streams while preserving previously acquired knowledge. However, current approaches face a critical limitation we term unstructured knowledge interference, a critical issue that arises when unconstrained parameter updates entangle discriminative representations across classes, severely contaminating the feature space and introducing significant transfer and bias risks. To address these challenges, we propose the Tree of Prompts (ToP), a novel hierarchical prompting framework that facilitates structured knowledge adaptation through multi-granular parameter regulation. ToP hierarchically integrates three synergistic components: (1) Stage-level prompts preserve historical knowledge by isolating task-specific parameters, thereby mitigating conflicts between incremental tasks; (2) Centroid-level prompts disentangle category semantics through learnable prototype calibration, sharpening decision boundaries in the feature space; and (3) Context-level prompts dynamically capture discriminative local features to suppress contamination from superficial similarities. Experimental results demonstrate that ToP markedly outperforms existing methods and provides a comprehensive and efficient solution for C-GCD. Yiqing Hao, Yangru Huang, Yi Jin 0001, Tao Wang 0011, Yidong Li, Yi-Gang Cen |
ACM Multimedia | 4 |
| 2025 | SALS: Sparse Attention in Latent Space for KV Cache CompressionabstractLarge Language Models (LLMs) capable of handling extended contexts are in high demand, yet their inference remains challenging due to substantial Key-Value (KV) cache size and high memory bandwidth requirements. Previous research has demonstrated that KV cache exhibits low-rank characteristics within the hidden dimension, suggesting the potential for effective compression. However, due to the widely adopted Rotary Position Embedding (RoPE) mechanism in modern LLMs, naive low‑-rank compression suffers severe accuracy degradation or creates a new speed bottleneck, as the low-rank cache must first be reconstructed in order to apply RoPE. In this paper, we introduce two key insights: first, the application of RoPE to the key vectors increases their variance, which in turn results in a higher rank; second, after the key vectors are transformed into the latent space, they largely maintain their representation across most layers. Based on these insights, we propose the Sparse Attention in Latent Space (SALS) framework. SALS projects the KV cache into a compact latent space via low-rank projection, and performs sparse token selection using RoPE-free query--key interactions in this space. By reconstructing only a small subset of important tokens, it avoids the overhead of full KV cache reconstruction. We comprehensively evaluate SALS on various tasks using two large-scale models: LLaMA2-7b-chat and Mistral-7b, and additionally verify its scalability on the RULER-128k benchmark with LLaMA3.1-8B-Instruct. Experimental results demonstrate that SALS achieves SOTA performance by maintaining competitive accuracy. Under different settings, SALS achieves 6.4-fold KV cache compression and 5.7-fold speed-up in the attention operator compared to FlashAttention2 on the 4K sequence. For the end-to-end throughput performance, we achieves 1.4-fold and 4.5-fold improvement compared to GPT-fast on 4k and 32K sequences, respectively. The source code will be publicly available in the future. Junlin Mu, Hantao Huang, Jihang Zhang, Minghui Yu, Tao Wang 0011, Yidong Li |
NeurIPS | 5 |
| 2025 | Open World Adaptive Pseudo Contrastive Learning for Generalized Category DiscoveryabstractIn this work, we investigate the challenging task of Generalized Category Discovery (GCD). Given datasets collected from open-world scenarios comprising both labeled and unlabeled images, GCD aims to classify all unlabeled images while simultaneously identifying unlabeled novel categories. The fundamental challenge in GCD tasks stems from inherent annotation discrepancies between seen and novel classes within the dataset. The lack of reliable label supervision for novel classes in unlabeled data leads to significant disparities in the model’s learning between old and novel classes, which is termed the bias risk. Recent advancements in GCD have employed the entropy maximization algorithm to alleviate the bias risk. However, they fail to provide debiased optimization for unlabeled data, leading to models that struggle with extracting discriminative features from such data. To address these challenges, we have created an Open-world pseudo-contrastive learning framework named OpcGCD. Our OpcGCD framework implements a dynamic category-wise threshold mechanism, which employs a parametric prototype classifiers to generate debiased pseudo-labels for unlabeled samples. To facilitate the learning of discriminative feature representations, our proposed OpcGCD employs debiased pseudo-labels in the formulation of a contrastive learning loss. Extensive evaluations conducted on multiple GCD benchmark datasets demonstrate the robustness and effectiveness of the approach. Yiqing Hao, Xu Wang 0053, Yi Jin 0001, Tao Wang 0011, Yidong Li, Shuoyan Liu, Chao Li 0026, Hui Yu 0001 |
SMC | 4 |
| 2025 | UNAGI: Unified neighbor-aware graph neural network for multi-view clustering
Zheming Xu, Congyan Lang, Liqian Liang, Tao Wang 0011, Yidong Li, Michael Kampffmeyer |
Neural Networks | 5 |
| 2025 | The Cascaded Forward algorithm for neural network training
Gongpei Zhao, Tao Wang 0011, Yi Jin 0001, Congyan Lang, Yidong Li, Haibin Ling |
Pattern Recognit. | 2 |
| 2025 | Learning Diversified Primitive Prompts for Compositional Zero-Shot Learning
Xinru Zhao, Congyan Lang, Tao Wang 0011, Yidong Li |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | M3-ReID: Unifying Multi-View, Granularity, and Modality for Video-Based Visible-Infrared Person Re-IdentificationabstractVideo-based visible-infrared person re-identification (VVI-ReID) task focuses on cross-modality retrieval of pedestrian videos, which are captured in visible and infrared modalities by non-overlapping cameras across diverse scenes, and holds significant value for security surveillance scenarios. The challenges of this task mainly stem from three issues: the difficulty of capturing comprehensive spatio-temporal cues, intra-class variations within video sequences, and inter-modality discrepancies between visible and infrared data. Existing methods mainly try to address the modality gap or focus on one of the other aspects, but rarely do they jointly consider these key factors. Motivated by these core challenges, we propose the M3-ReID (Multi-View & Granularity & Modality) method, a unified framework that simultaneously enhances spatio-temporal feature extraction, intra-class discrimination, and cross-modality consistency. Specifically, to capture diverse spatio-temporal patterns, we design a Multi-View Learning module that leverages different spatial and temporal-spatial perspectives to adaptively emphasize diverse key regions and motion cues. To enhance intra-class modeling of each identity, we introduce a Multi-Granularity Representation strategy that optimizes features across both fine-grained frame level and coarse-grained video level by minimizing mutual information among redundant frames while enhancing identity representations. Furthermore, to bridge the visible-infrared gap, we propose a Multi-Modality Alignment mechanism that explicitly aligns metric learning and cross-modality matching goals, transforming features into a unified embedding space with modality consistency and class discrimination. Extensive experiments on benchmark VVI-ReID datasets demonstrate the superiority of our proposed M3-ReID framework against existing methods. Tengfei Liang, Yi Jin 0001, Zhun Zhong, Xin Chen 0003, Xianjia Meng, Tao Wang 0011, Yidong Li |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2025 | Deep Probabilistic Graph MatchingabstractMost previous learning-based graph matching algorithms solve the quadratic assignment problem (QAP) by dropping one or more of the matching constraints and adopting a relaxed assignment solver to obtain sub-optimal correspondences. Such relaxation may actually weaken the original graph matching problem, and in turn hurt the matching performance. In this paper, we propose a deep learning-based graph matching framework that works for the original QAP without compromising on the matching constraints. In particular, we design an affinityassignment prediction network to jointly learn the pairwise affinity and estimate the node assignments, and we then develop a differentiable solver inspired by the probabilistic perspective of the pairwise affinities. Aiming to obtain better matching results, the probabilistic solver refines the estimated assignments in an iterative manner to impose both discrete and one-to-one matching constraints. The proposed method is trained in a supervised manner, evaluated on several benchmarks related to semantic keypoint corresponding, matching of social networks and pure QAP instances. In all experiment, it exhibits state-of-the-art matching performance on all benchmarks. Tao Wang 0011, Congyan Lang, Yidong Li, Haibin Ling |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2025 | LighTN: Light-Weight Transformer Network for Performance-Overhead Tradeoff in Point Cloud DownsamplingabstractDownsampling is a crucial task for processing large scale and/or dense point clouds with limited resources. Owing to the development of deep learning, approaches of task-oriented point cloud downsampling have significant performance gains in preserving geometric information. However, most downsamling methods are limited by the disordered and unstructured point cloud data, making it difficult to continually improve the performance. To address this issue, we propose a light-weight Transformer network (LighTN) for the task-oriented point cloud downsampling as an end-to-end solution. In LighTN, we design an energy-efficient and permutation invariant single-head self-correlation module to extract refined global geometric features. Moreover, we present a novel sampling loss function to guide LighTN to focus on critical point cloud regions with more uniform distributions and prominent point coverage. Extensive experiments on classification, registration, and reconstruction tasks demonstrate that LighTN can achieve the state-of-the-art performance-overhead tradeoff and high-quality qualitative results. Xu Wang 0053, Yi Jin 0001, Yi-Gang Cen, Tao Wang 0011, Bowen Tang 0001, Yidong Li |
IEEE Trans. Multim. | 4 |
| 2024 | Enhancing Multimedia Applications by Removing Dynamic Objects in Neural Radiance Fields
XianBen Yang, Tao Wang 0011, Yi Jin 0001, Congyan Lang, Yidong Li |
ACCV (10) | 2 |
| 2024 | SynSP: Synergy of Smoothness and Precision in Pose Sequences RefinementabstractPredicting human pose sequences via existing pose estimators often encounters various estimation errors. Motion refinement methods aim to optimize the predicted human pose sequences from pose estimators while ensuring minimal computational overhead and latency. Prior investigations have primarily concentrated on striking a balance between the two objectives, i.e., smoothness and precision, while optimizing the predicted pose sequences. However, it has come to our attention that the tension between these two objectives can provide additional quality cues about the predicted pose sequences. These cues, in turn, are able to aid the network in optimizing lower-quality poses. To leverage this quality information, we propose a motion refinement network, termed SynSP, to achieve a Synergy of Smoothness and Precision in the sequence refinement tasks. Moreover, SynSP can also address multi-view poses of one person simultaneously, fixing inaccuracies in predicted poses through heightened attention to similar poses from other views, thereby amplifying the resultant quality cues and overall performance. Compared with previous methods, SynSP benefits from both pose quality and multi-view information with a much shorter input sequence length, achieving state-of-the-art results among four challenging datasets involving 2D, 3D, and SMPL pose representations in both single-view and multi-view scenes. Github code: https://github.com/InvertedForest/SynSP. Tao Wang 0011, Lei Jin 0003, Zheng Wang 0007, Jianshu Li, Liang Li 0003, Fang Zhao 0006, Yu Cheng 0009, Li Yuan 0007, Junliang Xing, Jian Zhao 0006 |
CVPR | 1 |
| 2024 | Generated and Pseudo Content guided Prototype Refinement for Few-shot Point Cloud SegmentationabstractFew-shot 3D point cloud semantic segmentation aims to segment query point clouds with only a few annotated support point clouds. Existing prototype-based methods learn prototypes from the 3D support set to guide the segmentation of query point clouds. However, they encounter the challenge of low prototype quality due to constrained semantic information in the 3D support set and class information bias between support and query sets. To address these issues, in this paper, we propose a novel framework called Generated and Pseudo Content guided Prototype Refinement (GPCPR), which explicitly leverages LLM-generated content and reliable query context to enhance prototype quality. GPCPR achieves prototype refinement through two core components: LLM-driven Generated Content-guided Prototype Refinement (GCPR) and Pseudo Query Context-guided Prototype Refinement (PCPR). Specifically, GCPR integrates diverse and differentiated class descriptions generated by large language models to enrich prototypes with comprehensive semantic knowledge. PCPR further aggregates reliable class-specific pseudo-query context to mitigate class information bias and generate more suitable query-specific prototypes. Furthermore, we introduce a dual-distillation regularization term, enabling knowledge transfer between early-stage entities (prototypes or pseudo predictions) and their deeper counterparts to enhance refinement. Extensive experiments demonstrate the superiority of our method, surpassing the state-of-the-art methods by up to 12.10% and 13.75% mIoU on S3DIS and ScanNet, respectively. Congyan Lang, Tao Wang 0011, Yidong Li, Jun Liu 0036 |
NeurIPS | 4 |
| 2024 | DFA-GNN: Forward Learning of Graph Neural Networks by Direct Feedback AlignmentabstractGraph neural networks (GNNs) are recognized for their strong performance across various applications, with the backpropagation (BP) algorithm playing a central role in the development of most GNN models. However, despite its effectiveness, BP has limitations that challenge its biological plausibility and affect the efficiency, scalability and parallelism of training neural networks for graph-based tasks. While several non-backpropagation (non-BP) training algorithms, such as the direct feedback alignment (DFA), have been successfully applied to fully-connected and convolutional network components for handling Euclidean data, directly adapting these non-BP frameworks to manage non-Euclidean graph data in GNN models presents significant challenges. These challenges primarily arise from the violation of the independent and identically distributed (i.i.d.) assumption in graph data and the difficulty in accessing prediction errors for all samples (nodes) within the graph. To overcome these obstacles, in this paper we propose DFA-GNN, a novel forward learning framework tailored for GNNs with a case study of semi-supervised learning. The proposed method breaks the limitations of BP by using a dedicated forward training mechanism. Specifically, DFA-GNN extends the principles of DFA to adapt to graph data and unique architecture of GNNs, which incorporates the information of graph topology into the feedback links to accommodate the non-Euclidean characteristics of graph data. Additionally, for semi-supervised graph learning tasks, we developed a pseudo error generator that spreads residual errors from training data to create a pseudo error for each unlabeled node. These pseudo errors are then utilized to train GNNs using DFA. Extensive experiments on 10 public benchmarks reveal that our learning framework outperforms not only previous non-BP methods but also the standard BP methods, and it exhibits excellent robustness against various types of noise and attacks. Gongpei Zhao, Tao Wang 0011, Congyan Lang, Yi Jin 0001, Yidong Li, Haibin Ling |
NeurIPS | 2 |
| 2024 | Enhancing Point Cloud Sampling Quality with Dual-Branch Fusion NetworksabstractTask-oriented point cloud sampling methods have attracted considerable attention for their ability to adaptively select important point sets based on downstream tasks, achieving an excellent balance between data simplification and task performance. However, existing task-oriented sampling models, primarily based on single-branch designs, struggle to fully extract features from input point clouds that comprehensively reflect multi-dimensional key information, thus limiting their sampling performance. In this paper, we introduce a dual-branch sampling network, named DBS-NET, which conducts crucial point sampling from both the global and local importance perspectives separately before merging them, thereby preserving multi-dimensional key information of the input data during the sampling process. Qualitative and quantitative experimental results demonstrate the competitive performance of DBS-NET on the classification benchmark task. Yi Jin 0001, Xu Wang 0053, Mengxia Hu, Hui Yu 0001, Yidong Li, Tao Wang 0011, Songhe Feng, Congyan Lang |
SMC | 6 |
| 2024 | GLAN: A graph-based linear assignment network
Tao Wang 0011, Congyan Lang, Songhe Feng, Yi Jin 0001, Yidong Li |
Pattern Recognit. | 2 |
| 2024 | Bridging the Gap: Multi-Level Cross-Modality Joint Alignment for Visible-Infrared Person Re-IdentificationabstractVisible-Infrared person Re-IDentification (VI-ReID) is a challenging cross-modality image retrieval task that aims to match pedestrians’ images across visible and infrared cameras. To solve the modality gap, existing mainstream methods adopt a learning paradigm converting the image retrieval task into an image classification task with cross-entropy loss and auxiliary metric learning losses. These losses follow the strategy of adjusting the distribution of extracted embeddings to reduce the intra-class distance and increase the inter-class distance. However, such objectives do not precisely correspond to the final test setting of the retrieval task, resulting in a new gap at the optimization level. By rethinking these keys of VI-ReID, we propose a simple and effective method, the Multi-level Cross-modality Joint Alignment (MCJA), bridging both the modality and objective-level gap. For the former, we design the Visible-Infrared Modality Coordinator in the image space and propose the Modality Distribution Adapter in the feature space, effectively reducing modality discrepancy of the feature extraction process. For the latter, we introduce a new Cross-Modality Retrieval loss. It is the first work to constrain from the perspective of the ranking list in the VI-ReID, aligning with the goal of the testing stage. Moreover, to strengthen the robustness and cross-modality retrieval ability, we further introduce a Multi-Spectral Enhanced Ranking strategy for the testing phase. Based on the global feature only, our method outperforms existing methods by a large margin, achieving the remarkable rank-1 of 89.51% and mAP of 87.58% on the most challenging single-shot setting and all-search mode of the SYSU-MM01 dataset. Tengfei Liang, Yi Jin 0001, Wu Liu 0005, Tao Wang 0011, Songhe Feng, Yidong Li |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Linkage-Based Object Re-Identification via Graph LearningabstractObject Re-identification (Re-ID), which includes person Re-ID and vehicle Re-ID, is one of the core technologies of the intelligent transportation system. Existing supervised Re-ID studies mainly focus on discriminative feature learning (e.g., attention-based methods) or metric learning (e.g., triplet-loss-based methods) to obtain more accurate matches between the probe object and the positive gallery. However, they both pay less attention to global structure information (GSI) buried in the overall datasets. In this paper, we go beyond the traditional methods that are either unaware of or locally perceiving to GSI, and consider exploring the structural relationships among all the object instances of a dataset via a graph. Specifically, we construct a graph across the entire dataset, where each object instance is treated as a node and edges are assigned with the help of a classic algorithm like KNN. Seeing that a binary edge label can be used to predict whether its associated nodes belong to the same identity, we naturally formulate the problem of Re-ID as a new link prediction problem. Inspired by the superior capacity of capturing structure information of graph convolutional networks (GCN), a GCN-based global structure embedded network (GSE-Net) is proposed to take the graph as input and output a set of linkage likelihoods. During testing, we perform the evaluation according to the node features or estimated linkage likelihood via a graph where nodes include query and gallery images. Extensive experiments demonstrate that our proposed method outperforms the state-of-the-arts on both person and vehicle Re-ID benchmarks. Zhenxue Wang, Congyan Lang, Liqian Liang, Tao Wang 0011, Songhe Feng, Yidong Li |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | Discrete Listwise Content-aware RecommendationabstractTo perform online inference efficiently, hashing techniques, devoted to encoding model parameters as binary codes, play a key role in reducing the computational cost of content-aware recommendation (CAR), particularly on devices with limited computation resource. However, current hashing methods for CAR fail to align their learning objectives (e.g., squared loss) with the ranking-based metrics (e.g., Normalized Discounted Cumulative Gain (NDCG)), resulting in suboptimal recommendation accuracy. In this article, we propose a novel ranking-based CAR hashing method based on Factorization Machine (FM), called Discrete Listwise FM (DLFM), for fast and accurate recommendation. Concretely, our DLFM is to optimize NDCG in the Hamming space for preserving the listwise user-item relationships. We devise an efficient algorithm to resolve the challenging DLFM problem, which can directly learn binary parameters in a relaxed continuous solution space, without additional quantization. Particularly, our theoretical analysis shows that the optimal solution to the relaxed continuous optimization problem is approximately the same as that of the original discrete optimization problem. Through extensive experiments on two real-world datasets, we show that DLFM consistently outperforms state-of-the-art hashing-based recommendation techniques. Fangyuan Luo, Jun Wu 0007, Tao Wang 0011 |
ACM Trans. Knowl. Discov. Data | 3 |
| 2024 | Neighborhood Pattern Is Crucial for Graph Convolutional Networks Performing Node ClassificationabstractGraph convolutional networks (GCNs) are widely believed to perform well in the graph node classification task, and homophily assumption plays a core rule in the design of previous GCNs. However, some recent advances on this area have pointed out that homophily may not be a necessity for GCNs. For deeper analysis of the critical factor affecting the performance of GCNs, we first propose a metric, namely, neighborhood class consistency (NCC), to quantitatively characterize the neighborhood patterns of graph datasets. Experiments surprisingly illustrate that our NCC is a better indicator, in comparison to the widely used homophily metrics, to estimate GCN performance for node classification. Furthermore, we propose a topology augmentation graph convolutional network (TA-GCN) framework under the guidance of the NCC metric, which simultaneously learns an augmented graph topology with higher NCC score and a node classifier based on the augmented graph topology. Extensive experiments on six public benchmarks clearly show that the proposed TA-GCN derives ideal topology with higher NCC score given the original graph topology and raw features, and it achieves excellent performance for semi-supervised node classification in comparison to several state-of-the-art (SOTA) baseline algorithms. Gongpei Zhao, Tao Wang 0011, Yidong Li, Yi Jin 0001, Congyan Lang, Songhe Feng |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | P2ANet: A Large-Scale Benchmark for Dense Action Detection from Table Tennis Match Broadcasting VideosabstractWhile deep learning has been widely used for video analytics, such as video classification and action detection, dense action detection with fast-moving subjects from sports videos is still challenging. In this work, we release yet another sports video benchmark P 2 ANet for P ing P ong- A ction detection, which consists of 2,721 video clips collected from the broadcasting videos of professional table tennis matches in World Table Tennis Championships and Olympiads. We work with a crew of table tennis professionals and referees on a specially designed annotation toolbox to obtain fine-grained action labels (in 14 classes) for every ping-pong action that appeared in the dataset, and formulate two sets of action detection problems— action localization and action recognition . We evaluate a number of commonly seen action recognition (e.g., TSM, TSN, Video SwinTransformer, and Slowfast) and action localization models (e.g., BSN, BSN++, BMN, TCANet), using P 2 ANet for both problems, under various settings. These models can only achieve 48% area under the AR-AN curve for localization and 82% top-one accuracy for recognition since the ping-pong actions are dense with fast-moving subjects but broadcasting videos are with only 25 FPS. The results confirm that P 2 ANet is still a challenging task and can be used as a special benchmark for dense action detection from videos. We invite readers to examine our dataset by visiting the following link: https://github.com/Fred1991/P2ANET . Jiang Bian 0003, Xuhong Li 0002, Tao Wang 0011, Qingzhong Wang, Feixiang Lu, Dejing Dou, Haoyi Xiong |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | DecenterNet: Bottom-Up Human Pose Estimation Via Decentralized Pose RepresentationabstractMulti-person pose estimation in crowded scenes remains a very challenging task. This paper finds that most previous methods fail to estimate or group visible keypoints in crowded scenes rather than reasoning invisible keypoints. We thus categorize the crowded scenes into entanglement and occlusion based on the visibility of human parts and observe that entanglement is a significant problem in crowded scenes. With this observation, we propose DecenterNet, an end-to-end deep architecture to perform robust and efficient pose estimation in crowded scenes. Within DecenterNet, we introduce a decentralized pose representation that uses all visible keypoints as the root points to represent human poses, which is more robust in the entanglement area. We also propose a decoupled pose assessment mechanism, which introduces a location map to adaptively select optimal poses in the offset map. In addition, we have constructed a new dataset named SkatingPose, containing more entangled scenes. The proposed DecenterNet surpasses the best method on SkatingPose by 1.8 AP. Furthermore, DecenterNet obtains 71.2 AP and 71.4 AP on the COCO and CrowdPose datasets, respectively, demonstrating the superiority of our method. We will release our source code, trained models, and dataset to facilitate further studies in this research direction. Our code and dataset are available in https://github.com/InvertedForest/DecenterNet. Tao Wang 0011, Lei Jin 0003, Xiaojin Fan, Yu Cheng 0009, Yinglei Teng, Junliang Xing, Jian Zhao 0006 |
ACM Multimedia | 1 |
| 2023 | Depth guided feature selection for RGBD salient object detection
Zun Li 0001, Congyan Lang, Guanqin Li, Tao Wang 0011, Yidong Li |
Neurocomputing | 4 |
| 2023 | Object detection via inner-inter relational reasoning network
XiuTing You, Tao Wang 0011, Yidong Li |
Image Vis. Comput. | 3 |
| 2023 | Centroid-based graph matching networks for planar object tracking
Tao Wang 0011 |
Mach. Vis. Appl. | 3 |
| 2023 | Joint Graph Learning and Matching for Semantic Feature Correspondence
Tao Wang 0011, Yidong Li, Congyan Lang, Yi Jin 0001, Haibin Ling |
Pattern Recognit. | 2 |
| 2023 | Prior Knowledge Regularized Self-Representation Model for Partial Multilabel LearningabstractPartial multilabel learning (PML) aims to learn from training data, where each instance is associated with a set of candidate labels, among which only a part is correct. The common strategy to deal with such a problem is disambiguation, that is, identifying the ground-truth labels from the given candidate labels. However, the existing PML approaches always focus on leveraging the instance relationship to disambiguate the given noisy label space, while the potentially useful information in label space is not effectively explored. Meanwhile, the existence of noise and outliers in training data also makes the disambiguation operation less reliable, which inevitably decreases the robustness of the learned model. In this article, we propose a prior label knowledge regularized self-representation PML approach, called PAKS, where the self-representation scheme and prior label knowledge are jointly incorporated into a unified framework. Specifically, we introduce a self-representation model with a low-rank constraint, which aims to learn the subspace representations of distinct instances and explore the high-order underlying correlation among different instances. Meanwhile, we incorporate prior label knowledge into the above self-representation model, where the prior label knowledge is regarded as the complement of features to obtain an accurate self-representation matrix. The core of PAKS is to take advantage of the data membership preference, which is derived from the prior label knowledge, to purify the discovered membership of the data and accordingly obtain more representative feature subspace for model induction. Enormous experiments on both synthetic and real-world datasets show that our proposed approach can achieve superior or comparable performance to state-of-the-art approaches. Gengyu Lyu, Songhe Feng, Yi Jin 0001, Tao Wang 0011, Congyan Lang, Yidong Li |
IEEE Trans. Cybern. | 4 |
| 2023 | Beyond Word Embeddings: Heterogeneous Prior Knowledge Driven Multi-Label Image ClassificationabstractMulti-Label Image Classification (MLIC) is a fundamental yet challenging task which aims to recognize multiple labels from given images. The key to solve MLIC lies in how to accurately model the correlation between labels. Recent studies often adopt Graph Convolutional Network (GCN) to model label dependencies with word embeddings as prior knowledge. However, classical word embeddings typically contain redundant information due to the imperfect distributional hypothesis it relies on, which may degrade model generalizability. To tackle this problem, we propose a novel deep learning framework termedVisual-Semantic basedGraphConvolutionalNetwork (VSGCN), which alleviates the negative impact of redundant information by utilizing heterogeneous sources of prior knowledge. Specifically, we construct both visual prototype and semantic prototype for each label as heterogeneous prior label representations, which are further mapped to multi-label classifiers via two Multi-Head GCNs separately. The Multi-Head GCN mechanism proposed in this paper aims to guide the information propagation between prototypes for each label, which constructs multiple correlation graphs to simultaneously model the label correlation in different subspaces. Notably, we alleviate the negative influence of needless information by decreasing the inconsistency of predictions that come from visual space and semantic space. Extensive experiments conducted on various multi-label image datasets demonstrate the superiority of our proposed method. Songhe Feng, Gengyu Lyu, Tao Wang 0011, Congyan Lang |
IEEE Trans. Multim. | 4 |
| 2023 | SSR-Net: A Spatial Structural Relation Network for Vehicle Re-identificationabstractVehicle re-identification (Re-ID) represents the task aiming to identify the same vehicle from images captured by different cameras. Recent years have seen various feature learning-based approaches merely focusing on feature representations including global features or local features to obtain more subtle details to identify highly similar vehicles. However, few such methods consider the spatial geometrical structure relationship among local regions or between the global and local regions. By contrast, in this study, we propose a Spatial Structural Relation Network (SSR-Net) that explores the above-mentioned two kinds of relations simultaneously to learn more discriminative features by modeling the spatial structure information and global context information. In this article, we propose to adopt a Graph Convolution Network (GCN), for modeling spatial structural relationships among characteristic features. The GCN model aggregating the local and global features is shown to be more discriminative and robust to several car image transformations. To improve the performance of our proposed network, we jointly combine the classification loss with metric learning loss. Extensive experiments conducted on the public VehicleID and VeRi-776 datasets validate the effectiveness of our approach in comparison with recent works. Zheming Xu, Congyan Lang, Songhe Feng, Tao Wang 0011, Adrian G. Bors, Hongzhe Liu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2022 | Camera-Aware Style Separation and Contrastive Learning for Unsupervised Person Re-IdentificationabstractUnsupervised person re-identification (ReID) is a challenging task without data annotation to guide discriminative learning. Existing methods attempt to solve this problem by clustering extracted embeddings to generate pseudo labels. However, most methods ignore the intra-class gap caused by camera style variance, and some methods are relatively complex and indirect although they try to solve the negative impact of the camera style on feature distribution. To solve this problem, we propose a camera-aware style separation and contrastive learning method (CA-UReID), which directly separates camera styles in the feature space with the designed camera-aware attention module. It can explicitly divide the learnable feature into camera-specific and camera-agnostic parts, reducing the influence of different cameras. Moreover, to further narrow the gap across cameras, we design a camera-aware contrastive center loss to learn more discriminative embeddings for each identity. Extensive experiments demonstrate the superiority of our method over the state-of-the-art methods on the unsupervised person ReID task. Tengfei Liang, Yi Jin 0001, Tao Wang 0011, Yidong Li |
ICME | 4 |
| 2022 | Discrete Listwise Personalized Ranking for Fast Top-N Recommendation with Implicit FeedbackabstractWe address the efficiency problem of personalized ranking from implicit feedback by hashing users and items with binary codes, so that top-N recommendation can be fast executed in a Hamming space by bit operations. However, current hashing methods for top-N recommendation fail to align their learning objectives (such as pointwise or pairwise loss) with the benchmark metrics for ranking quality (e.g. Average Precision, AP), resulting in sub-optimal accuracy. To this end, we propose a Discrete Listwise Personalized Ranking (DLPR) model that optimizes AP under discrete constraints for fast and accurate top-N recommendation. To resolve the challenging DLPR problem, we devise an efficient algorithm that can directly learn binary codes in a relaxed continuous solution space. Specifically, theoretical analysis shows that the optimal solution to the relaxed continuous optimization problem is exactly the same as that of the original discrete DLPR problem. Through extensive experiments on two real-world datasets, we show that DLPR consistently surpasses state-of-the-art hashing methods for top-N recommendation. Fangyuan Luo, Jun Wu 0007, Tao Wang 0011 |
IJCAI | 3 |
| 2022 | Keypoint-Guided Modality-Invariant Discriminative Learning for Visible-Infrared Person Re-identificationabstractThe visible-infrared person re-identification (VI-ReID) task aims to retrieve images of pedestrians across cameras with different modalities. In this task, the major challenges arise from two aspects: intra-class variations among images of the same identity, and cross-modality discrepancies between visible and infrared images. Existing methods mainly focus on the latter, attempting to alleviate the impact of modality discrepancy, which ignore the former issue of identity variations and achieve limited discrimination. To address both aspects, we propose a Keypoint-guided Modality-invariant Discriminative Learning (KMDL) method, which can simultaneously adapt to intra-ID variations and bridge the cross-modality gap. By introducing human keypoints, our method makes further exploration in the image space, feature space and loss constraints to solve the above issues. Specifically, considering the modality discrepancy in original images, we first design a Hue Jitter Augmentation (HJA) strategy, introducing the hue disturbance to alleviate color dependence in the input stage. To obtain discriminative fine-grained representation for retrieval, we design the Global-Keypoint Graph Module (GKGM) in feature space, which can directly extract keypoint-aligned features and mine relationships within global and keypoint embeddings. Based on these semantic local embeddings, we further propose the Keypoint-Aware Center (KAC) loss that can effectively adjust the feature distribution under the supervision of ID and keypoint to learn discriminative representation for the matching. Extensive experiments on SYSU-MM01 and RegDB datasets demonstrate the effectiveness of our KMDL method. Tengfei Liang, Yi Jin 0001, Wu Liu 0005, Songhe Feng, Tao Wang 0011, Yidong Li |
ACM Multimedia | 5 |
| 2022 | Object representation enhancement for self-supervised colocalizationabstractSelf-supervised colocalization is to localize common objects in the data set containing only one superclass without using human-annotated labels. Existing methods achieve impressive results by employing self-supervised pretext learning. However, a common limitation still exists. They either tend to overextend activations to the background, or they tend to activate the most discriminative object part. To alleviate this problem, we propose an object representation enhancement model to weaken background distraction and to mine complementary object regions during the object representation learning. Specifically, we first propose an Object-aware Representation Enhancement (ORE) module to estimate an object mask for each input image, guiding the model to disregard the background content and focus on the foreground object. The ORE module and the subsequent self-supervised learning can mutually reinforce each other. Then we propose a Masked Self-supervised Learning branch and design a masked attention consistency objective to induce the model to activate complementary parts of the object effectively. Extensive experiments on four fine-grained data sets demonstrate the superiority of the proposed model. Yidong Li, Yi Jin 0001, Tao Wang 0011 |
Int. J. Intell. Syst. | 4 |
| 2022 | Pedestrian attribute recognition based on attribute correlation
Ruijie Zhao 0007, Congyan Lang, Zun Li 0001, Liqian Liang, Songhe Feng, Tao Wang 0011 |
Multim. Syst. | 7 |
| 2022 | Object detection by crossing relational reasoning based on graph neural network
XiuTing You, Tao Wang 0011, Songhe Feng, Congyan Lang |
Mach. Vis. Appl. | 3 |
| 2022 | A Self-Paced Regularization Framework for Partial-Label LearningabstractPartial-label learning (PLL) aims to solve the problem where each training instance is associated with a set of candidate labels, one of which is the correct label. Most PLL algorithms try to disambiguate the candidate label set, by either simply treating each candidate label equally or iteratively identifying the true label. Nonetheless, existing algorithms usually treat all labels and instances equally, and the complexities of both labels and instances are not taken into consideration during the learning stage. Inspired by the successful application of a self-paced learning strategy in the machine-learning field, we integrate the self-paced regime into the PLL framework and propose a novel self-paced PLL (SP-PLL) algorithm, which could control the learning process to alleviate the problem by ranking the priorities of the training examples together with their candidate labels during each learning iteration. Extensive experiments and comparisons with other baseline methods demonstrate the effectiveness and robustness of the proposed method. Gengyu Lyu, Songhe Feng, Tao Wang 0011, Congyan Lang |
IEEE Trans. Cybern. | 3 |
| 2022 | Weakly Supervised Video Object Segmentation via Dual-attention Cross-branch FusionabstractRecently, concerning the challenge of collecting large-scale explicitly annotated videos, weakly supervised video object segmentation (WSVOS) using video tags has attracted much attention. Existing WSVOS approaches follow a general pipeline including two phases, i.e., a pseudo masks generation phase and a refinement phase. To explore the intrinsic property and correlation buried in the video frames, most of them focus on the later phase by introducing optical flow as temporal information to provide more supervision. However, these optical flow-based studies are greatly affected by illumination and distortion and lack consideration of the discriminative capacity of multi-level deep features. In this article, with the goal of capturing more effective temporal information and investigating a temporal information fusion strategy accordingly, we propose a unified WSVOS model by adopting a two-branch architecture with a multi-level cross-branch fusion strategy, named as dual-attention cross-branch fusion network (DACF-Net). Concretely, the two branches of DACF-Net, i.e., a temporal prediction subnetwork (TPN) and a spatial segmentation subnetwork (SSN), are used for extracting temporal information and generating predicted segmentation masks, respectively. To perform the cross-branch fusion between TPN and SSN, we propose a dual-attention fusion module that can be plugged into the SSN flexibly. We also pose a cross-frame coherence loss (CFCL) to achieve smooth segmentation results by exploiting the coherence of masks produced by TPN and SSN. Extensive experiments demonstrate the effectiveness of proposed approach compared with the state-of-the-arts on two challenging datasets, i.e., Davis-2016 and YouTube-Objects. Congyan Lang, Liqian Liang, Songhe Feng, Tao Wang 0011, Shidi Chen |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2022 | VARID: Viewpoint-Aware Re-IDentification of Vehicle Based on Triplet LossabstractWith the increasing prevalence of intelligent traffic control and monitoring, research on vehicle re-identification (Re-ID) draws substantial attention in recent years. Different from other cross-view searching tasks such as person Re-ID, the vehicle Re-ID problem is more challenging and unpredictable as viewpoint variations can greatly affect the appearance of vehicles. Existing studies mainly focus on extracting global features based on visual appearance to represent the identity of the target vehicle, while the impact of viewpoint variation is rarely considered. In this paper, we take the view information into account to boost vehicle Re-ID, and introduce latent view labels by clustering and incorporates view information into deep metric learning to tackle the challenge. We also develop a stricter center constraint to further improve the intra-class compactness of feature space. Moreover, we adopt an orthogonal regularization to increase the separability between different vehicles. VARID achieves 79.3% mAP on VeRi-776 and 88.5% mAP on VehicleID which surpasses state-of-the-arts a lot. More comprehensive experimental analyses and evaluations on four benchmarks demonstrate that the proposed method outperforms significantly state-of-the-arts methods. Yidong Li, Yi Jin 0001, Tao Wang 0011, Weipeng Lin |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Seeing Crucial Parts: Vehicle Model Verification via a Discriminative Representation ModelabstractWidely used surveillance cameras have promoted large amounts of street scene data, which contains one important but long-neglected object: the vehicle. Here we focus on the challenging problem of vehicle model verification. Most previous works usually employ global features (e.g., fully connected features) to further perform vehicle-level deep metric learning (e.g., triplet-based network). However, we argue that it is noteworthy to investigate the distinctiveness of local features and consider vehicle-part-level metric learning by reducing the intra-class variance as much as possible. In this article, we introduce a simple yet powerful deep model—the enforced intra-class alignment network (EIA-Net)—which can learn a more discriminative image representation by localizing key vehicle parts and jointly incorporating two distance metrics: vehicle-level embedding and vehicle-part-sensitive embedding. For learning features, we propose an effective feature extraction module that is composed of two components: the regional proposal network (RPN)-based network and part-based CNN. The RPN is used to define key vehicle regions and aggregate local features on these regions, whereas part-based CNN offers supplementary global features for the RPN-based network. The fusion features learned by feature extraction module are cast into the deep metric learning module. Especially, we derived an enforced intra-class alignment loss by re-utilizing key vehicle part information to enhance reducing intra-class variance. Furthermore, we modify the coupled cluster loss to model the vehicle-level embedding by enlarging the inter-class variance while shortening intra-class variance. Extensive experiments over benchmark datasets VehicleID and CompCars have shown that the proposed EIA-Net significantly outperforms the state-of-the-art approaches for vehicle model verification. Furthermore, we also conduct comprehensive experiments on vehicle re-identification datasets (i.e., VehicleID and VeRi776) to validate the generalization ability effectiveness of our proposed method. Liqian Liang, Congyan Lang, Zun Li 0001, Jian Zhao 0006, Tao Wang 0011, Songhe Feng |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2021 | Deep spatio-frequency saliency detection
Zun Li 0001, Congyan Lang, Tao Wang 0011, Yidong Li, Jiashi Feng |
Neurocomputing | 3 |
| 2021 | Entropy-aware self-training for graph convolutional networks
Gongpei Zhao, Tao Wang 0011, Yidong Li, Yi Jin 0001, Congyan Lang |
Neurocomputing | 2 |
| 2021 | Text to photo-realistic image synthesis via chained deep recurrent generative adversarial network
Congyan Lang, Songhe Feng, Tao Wang 0011, Yi Jin 0001, Yidong Li |
J. Vis. Commun. Image Represent. | 4 |
| 2021 | Fine-Grained Semantic Image Synthesis with Object-Attention Generative Adversarial NetworkabstractSemantic image synthesis is a new rising and challenging vision problem accompanied by the recent promising advances in generative adversarial networks. The existing semantic image synthesis methods only consider the global information provided by the semantic segmentation mask, such as class label, global layout, and location, so the generative models cannot capture the rich local fine-grained information of the images (e.g., object structure, contour, and texture). To address this issue, we adopt a multi-scale feature fusion algorithm to refine the generated images by learning the fine-grained information of the local objects. We propose OA-GAN, a novel object-attention generative adversarial network that allows attention-driven, multi-fusion refinement for fine-grained semantic image synthesis. Specifically, the proposed model first generates multi-scale global image features and local object features, respectively, then the local object features are fused into the global image features to improve the correlation between the local and the global. In the process of feature fusion, the global image features and the local object features are fused through the channel-spatial-wise fusion block to learn ‘what’ and ‘where’ to attend in the channel and spatial axes, respectively. The fused features are used to construct correlation filters to obtain feature response maps to determine the locations, contours, and textures of the objects. Extensive quantitative and qualitative experiments on COCO-Stuff, ADE20K and Cityscapes datasets demonstrate that our OA-GAN significantly outperforms the state-of-the-art methods. Congyan Lang, Liqian Liang, Songhe Feng, Tao Wang 0011, Yutong Gao 0001 |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2021 | GM-PLL: Graph Matching Based Partial Label LearningabstractPartial Label Learning (PLL) aims to learn from the data where each training example is associated with a set of candidate labels, among which only one is correct. The key to deal with such problem is to disambiguate the candidate label sets and obtain the correct assignments between instances and their candidate labels. In this paper, we interpret such assignments as instance-to-label matchings, and reformulate the task of PLL as a matching selection problem. To model such problem, we propose a novel Graph Matching based Partial Label Learning (GM-PLL) framework, where Graph Matching (GM) scheme is incorporated owing to its excellent capability of exploiting the instance and label relationship. Meanwhile, since conventional one-to-one GM algorithm does not satisfy the constraint of PLL problem that multiple instances may correspond to the same label, we extend a traditional one-to-one probabilistic matching algorithm to the many-to-one constraint, and make the proposed framework accommodate to the PLL problem. Moreover, we also propose a relaxed matching prediction model, which can improve the prediction accuracy via GM strategy. Extensive experiments on both artificial and real-world data sets demonstrate that the proposed method can achieve superior or comparable performance against the state-of-the-art methods. Gengyu Lyu, Songhe Feng, Tao Wang 0011, Congyan Lang, Yidong Li |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2020 | Learning Combinatorial Solver for Graph MatchingabstractLearning-based approaches to graph matching have been developed and explored for more than a decade, have grown rapidly in scope and popularity in recent years. However, previous learning-based algorithms, with or without deep learning strategy, mainly focus on the learning of node and/or edge affinities generation, and pay less attention on the learning of the combinatorial solver. In this paper we propose a fully trainable framework for graph matching, in which learning of affinities and solving for combinatorial optimization are not explicitly separated as in many previous arts. We firstly convert the problem of building node correspondences between two input graphs to the problem of selecting reliable nodes from a constructed assignment graph. Subsequently, the graph network block module is adopted to perform computation on the graph to form structured representations for each node. It finally predicts a label for each node that is used for node classification, and the training is performed under the supervision of both permutation differences and the one-to-one matching constraints. The proposed method is evaluated on four public benchmarks in comparison with several state-of-the-art algorithms, and the experimental results illustrate its excellent performance. Tao Wang 0011, Yidong Li, Yi Jin 0001, Xiaohui Hou, Haibin Ling |
CVPR | 1 |
| 2020 | Attentive Generative Adversarial Network To Bridge Multi-Domain Gap For Image SynthesisabstractDespite the significant progress on text-to-image synthesis, automatically generating realistic images remains a challenging task since the location and specific shape of object are not given in the text descriptions. To address these problems, we propose a novel attentive generative adversarial network with contextual loss (AGAN-CL) algorithm. More specifically, the generative network consists of two sub-networks: a contextual network for generating image contours, and a cycle transformation autoencoder for converting contours to realistic images. Our core idea is the injection of image contours into the generative network, which is the most critical part of our network, since it will guide the whole generative network to focus on object regions. In addition, we also apply contextual loss and cycle-consistent loss to bridge multi-domain gap. Comprehensive results on several challenging datasets demonstrate the advantage of the proposed method over the leading approaches, regarding both visual fidelity and alignment with input descriptions. Congyan Lang, Liqian Liang, Gengyu Lyu, Songhe Feng, Tao Wang 0011 |
ICME | 6 |
| 2020 | End-to-End Text-to-Image Synthesis with Spatial ConstrainsabstractAlthough the performance of automatically generating high-resolution realistic images from text descriptions has been significantly boosted, many challenging issues in image synthesis have not been fully investigated, due to shapes variations, viewpoint changes, pose changes, and the relations of multiple objects. In this article, we propose a novel end-to-end approach for text-to-image synthesis with spatial constraints by mining object spatial location and shape information. Instead of learning a hierarchical mapping from text to image, our algorithm directly generates multi-object fine-grained images through the guidance of the generated semantic layouts. By fusing text semantic and spatial information into a synthesis module and jointly fine-tuning them with multi-scale semantic layouts generated, the proposed networks show impressive performance in text-to-image synthesis for complex scenes. We evaluate our method both on single-object CUB dataset and multi-object MS-COCO dataset. Comprehensive experimental results demonstrate that our method significantly outperforms the state-of-the-art approaches consistently across different evaluation metrics. Congyan Lang, Liqian Liang, Songhe Feng, Tao Wang 0011, Yutong Gao 0001 |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2020 | Learning part-alignment feature for person re-identification with spatial-temporal-based re-ranking method
Yi Jin 0001, Yidong Li, Congyan Lang, Songhe Feng, Tao Wang 0011 |
World Wide Web | 6 |
| 2019 | Partial Multi-Label Learning by Low-Rank and Sparse DecompositionabstractMulti-Label Learning (MLL) aims to learn from the training data where each example is represented by a single instance while associated with a set of candidate labels. Most existing MLL methods are typically designed to handle the problem of missing labels. However, in many real-world scenarios, the labeling information for multi-label data is always redundant , which can not be solved by classical MLL methods, thus a novel Partial Multi-label Learning (PML) framework is proposed to cope with such problem, i.e. removing the the noisy labels from the multi-label sets. In this paper, in order to further improve the denoising capability of PML framework, we utilize the low-rank and sparse decomposition scheme and propose a novel Partial Multi-label Learning by Low-Rank and Sparse decomposition (PML-LRS) approach. Specifically, we first reformulate the observed label set into a label matrix, and then decompose it into a groundtruth label matrix and an irrelevant label matrix, where the former is constrained to be low rank and the latter is assumed to be sparse. Next, we utilize the feature mapping matrix to explore the label correlations and meanwhile constrain the feature mapping matrix to be low rank to prevent the proposed method from being overfitting. Finally, we obtain the ground-truth labels via minimizing the label loss, where the Augmented Lagrange Multiplier (ALM) algorithm is incorporated to solve the optimization problem. Enormous experimental results demonstrate that PML-LRS can achieve superior or competitive performance against other state-of-the-art methods. Songhe Feng, Tao Wang 0011, Congyan Lang, Yi Jin 0001 |
AAAI | 3 |
| 2019 | Deformable Surface Tracking by Graph MatchingabstractThis paper addresses the problem of deformable surface tracking from monocular images. Specifically, we propose a graph-based approach that effectively explores the structure information of the surface to enhance tracking performance. Our approach solves simultaneously for feature correspondence, outlier rejection and shape reconstruction by optimizing a single objective function, which is defined by means of pairwise projection errors between graph structures instead of unary projection errors between matched points. Furthermore, an efficient matching algorithm is developed based on soft matching relaxation. For evaluation, our approach is extensively compared to state-of-the-art algorithms on a standard dataset of occluded surfaces, as well as a newly compiled dataset of different surfaces with rich, weak or repetitive texture. Experimental results reveal that our approach achieves robust tracking results for surfaces with different types of texture, and outperforms other algorithms in both accuracy and efficiency. Tao Wang 0011, Haibin Ling, Congyan Lang, Songhe Feng, Xiaohui Hou |
ICCV | 1 |
| 2019 | Recurrent convolutional network for video-based smoke detection
Mengxia Yin, Congyan Lang, Zun Li 0001, Songhe Feng, Tao Wang 0011 |
Multim. Tools Appl. | 5 |
| 2019 | Co-saliency Detection with Graph MatchingabstractRecently, co-saliency detection, which aims to automatically discover common and salient objects appeared in several relevant images, has attracted increased interest in the computer vision community. In this article, we present a novel graph-matching based model for co-saliency detection in image pairs. A solution of graph matching is proposed to integrate the visual appearance, saliency coherence, and spatial structural continuity for detecting co-saliency collaboratively. Since the saliency and the visual similarity have been seamlessly integrated, such a joint inference schema is able to produce more accurate and reliable results. More concretely, the proposed model first computes the intra-saliency for each image by aggregating multiple saliency cues. The common and salient regions across multiple images are thus discovered via a graph matching procedure. Then, a graph reconstruction scheme is proposed to refine the intra-saliency iteratively. Compared to existing co-saliency detection methods that only utilize visual appearance cues, our proposed model can effectively exploit both visual appearance and structure information to better guide co-saliency detection. Extensive experiments on several challenging image pair databases demonstrate that our model outperforms state-of-the-art baselines significantly. Zun Li 0001, Congyan Lang, Jiashi Feng, Yidong Li, Tao Wang 0011, Songhe Feng |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2018 | Constrained Confidence Matching for Planar Object TrackingabstractTracking planar objects has a wide range of applications in robotics. Conventional template tracking algorithms, however, often fail to observe fast object motion or drift significantly after a period of time, due to drastic object appearance change. To address such challenges, we propose a novel constrained confidence matching algorithm for motion estimation and a robust Kalman filter for template updating. Integrated with an accurate occlusion detector, our approach achieves accurate motion estimation in presence of partial occlusion, by excluding occluded pixels from computation of motion parameters. Furthermore, the proposed Kalman filter employs a novel control-input model to handle the object appearance change, which brings our tracker high robustness against sudden illumination change and heavy motion blur. For evaluation, we compare the proposed tracker with several state-of-the-art planar object trackers on two public benchmark datasets. Experimental results show that our algorithm achieves robust tracking results against various environmental variations, and outperforms baseline algorithms remarkably on both datasets. Tao Wang 0011, Haibin Ling, Congyan Lang, Songhe Feng, Yi Jin 0001, Yidong Li |
ICRA | 1 |
| 2018 | Hierarchical Discriminant Feature Learning for Heterogeneous Face RecognitionabstractHeterogeneous Face Recognition (HFR) refers to the problem of recognizing faces across different visual domains and has attached great attention owing to its tremendous potential benefits in practical applications. In this paper, a novel feature learning approach named hierarchical discriminant feature learning (HDFL) has been proposed for HFR. Different from traditional feature learning based HFR approaches, the proposed HDFL aims to learn the most discriminative information via a two-layer hierarchical boosting network (HBN), where the hierarchical discriminative information can be exploited in the learned features and the appearance difference can be effectively reduced, simultaneously. Extensive experiments on three different heterogeneous face databases demonstrate that our approach consistently outperforms the state-of-the-art methods. Yidong Li, Yi Jin 0001, Congyan Lang, Songhe Feng, Tao Wang 0011 |
VCIP | 6 |
| 2018 | Saliency ranker: A new salient object detection method
Zun Li 0001, Congyan Lang, Songhe Feng, Tao Wang 0011 |
J. Vis. Commun. Image Represent. | 4 |
| 2018 | A novel hypergraph matching algorithm based on tensor refining
Tao Wang 0011, Congyan Lang, Songhe Feng, Yi Jin 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2018 | Gracker: A Graph-Based Planar Object TrackerabstractMatching-based algorithms have been commonly used in planar object tracking. They often model a planar object as a set of keypoints, and then find correspondences between keypoint sets via descriptor matching. In previous work, unary constraints on appearances or locations are usually used to guide the matching. However, these approaches rarely utilize structure information of the object, and are thus suffering from various perturbation factors. In this paper, we proposed a graph-based tracker, named Gracker, which is able to fully explore the structure information of the object to enhance tracking performance. We model a planar object as a graph, instead of a simple collection of keypoints, to represent its structure. Then, we reformulate tracking as a sequential graph matching process, which establishes keypoint correspondence in a geometric graph matching manner. For evaluation, we compare the proposed Gracker with state-of-the-art planar object trackers on three benchmark datasets: two public ones and a newly collected one. Experimental results show that Gracker achieves robust tracking results against various environmental variations, and outperforms other algorithms in general on the datasets. Tao Wang 0011, Haibin Ling |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Graph Matching with Adaptive and Branching Path FollowingabstractGraph matching aims at establishing correspondences between graph elements, and is widely used in many computer vision tasks. Among recently proposed graph matching algorithms, those utilizing the path following strategy have attracted special research attentions due to their exhibition of state-of-the-art performances. However, the paths computed in these algorithms often contain singular points, which could hurt the matching performance if not dealt properly. To deal with this issue, we propose a novel path following strategy, named branching path following (BPF), to improve graph matching accuracy. In particular, we first propose a singular point detector by solving a KKT system, and then design a branch switching method to seek for better paths at singular points. Moreover, to reduce the computational burden of the BPF strategy, an adaptive path estimation (APE) strategy is integrated into BPF to accelerate the convergence of searching along each path. A new graph matching algorithm named ABPF-G is developed by applying APE and BPF to a recently proposed path following algorithm named GNCCP (Liu & Qiao 2014). Experimental results reveal how our approach consistently outperforms state-of-the-art algorithms for graph matching on five public benchmark datasets. Tao Wang 0011, Haibin Ling, Congyan Lang, Songhe Feng |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2017 | Multiple path exploration for graph matching
Congyan Lang, Tao Wang 0011 |
Mach. Vis. Appl. | 3 |
| 2017 | Human Facial Age Estimation by Cost-Sensitive Label Ranking and Trace Norm RegularizationabstractHuman facial age estimation has attracted much attention due to its potential applications in forensics, security, and biometrics. In contrast to existing approaches that cast facial age estimation as either a multiclass classification or regression problem, in this work, we propose a novel approach that combines the strength of cost-sensitive label ranking methods with the power of low-rank matrix recovery theories. Instead of having to make a binary decision for each age label, our approach ranks age labels in a descending order in terms of their predicted relevance to the given facial image. In addition, the proposed approach aggregates the linear prediction functions for different ages into a matrix, and introduces the matrix trace norm regularization to explicitly capture the correlations among different age labels and control the model complexity as well. Furthermore, motivated by nonlinear generalization performance of kernel methods, we extend the trace norm regularization from a finite dimensional space to an infinite dimensional space. We also provide theoretical analysis on the efficiency of the proposed kernelized trace normalization, which guarantees the feasibility of the proposed method for solving large-scale prediction problems. Comprehensive experiments on multiple well-known facial image datasets demonstrate the effectiveness of the proposed framework for age estimation compared to the state-of-the-arts. Songhe Feng, Congyan Lang, Jiashi Feng, Tao Wang 0011, Jiebo Luo 0001 |
IEEE Trans. Multim. | 4 |
| 2016 | Path Following with Adaptive Path Estimation for Graph MatchingabstractGraph matching plays an important role in many fields in computer vision. It is a well-known general NP-hard problem and has been investigated for decades. Among the large amount of algorithms for graph matching, the algorithms utilizing the path following strategy exhibited state-of-art performances. However, the main drawback of this category of algorithms lies in their high computational burden. In this paper, we propose a novel path following strategy for graph matching aiming to improve its computation efficiency. We first propose a path estimation method to reduce the computational cost at each iteration, and subsequently a method of adaptive step length to accelerate the convergence. The proposed approach is able to be integrated into all the algorithms that utilize the path following strategy. To validate our approach, we compare our approach with several recently proposed graph matching algorithms on three benchmark image datasets. Experimental results show that, our approach improves significantly the computation efficiency of the original algorithms, and offers similar or better matching results. Tao Wang 0011, Haibin Ling |
AAAI | 1 |
| 2016 | Branching Path Following for Graph Matching
Tao Wang 0011, Haibin Ling, Congyan Lang, Jun Wu 0007 |
ECCV (2) | 1 |
| 2016 | From sample selection to model update: A robust online visual tracking algorithm against drifting
Zhu Teng, Tao Wang 0011, Feng Liu 0061, Dong-Joong Kang, Congyan Lang, Songhe Feng |
Neurocomputing | 2 |
| 2016 | Symmetry-aware graph matching
Tao Wang 0011, Haibin Ling, Congyan Lang, Songhe Feng |
Pattern Recognit. | 1 |
| 2015 | An error-tolerant approximate matching algorithm for labeled combinatorial maps
Tao Wang 0011, Congyan Lang, Songhe Feng |
Neurocomputing | 1 |
| 2014 | Practical Anonymization for Protecting Privacy in Combinatorial MapsabstractCombinatorial Map (CM) is becoming increasingly popular due to its power in modeling topological structures with subdivided objects, which is widely used in the fields of social network, computer vision, social media and so on. However, due to its specific structural properties, an unprotected release of a combinatorial map may cause the identity disclosure problem, which is a major privacy breach revealing the identification of entities with certain background knowledge known by an adversary. In this paper, we discuss the privacy preserving problem in publishing private combinatorial maps. We first formalize a specific anonymizing model to deal with dart-related attacks, and discuss an efficient metric to quantify information loss incurred in the perturbation. Then we propose an efficient method for the dart anonymization problem to prevent a CM from the attack. Our approaches are efficient and practical, and have been validated by extensive experiments on two sets of synthetic data. Dandan Chu, Yidong Li, Tao Wang 0011, Hong Shen 0001 |
PDCAT | 3 |