EDBT 2026 Demo / reviewers in the wild / expert
Hua Yang 0001
dblp:11/4973-1
· DBLP profile ↗
81ranked-venue papers
4as first author
38since 2021 · last 2026
0000-0002-0417-234XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 63 · 2 first-author · 28 since 2021Artificial intelligence and machine learning · 26 · 2 first-author · 16 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OneLIP: Unlocking and Improving Long-Text Representations of CLIP via One-Stage AdaptationabstractContrastive Language-Image Pretraining (CLIP) has demonstrated impressive generalization on vision-language tasks by aligning images and short texts. However, its inherent 77-token length limits the capacity of capturing complex semantics in long captions. Existing long-text adaptations for CLIP typically rely on either multi-stage training or truncation-based alignment, both inevitably resulting in semantic degradation and cumbersome tuning. Therefore, we propose OneLIP, a unified framework that extends CLIP to understand long captions within a single training stage, eliminating the need for brittle truncation or multi-stage pipelines. OneLIP addresses semantic degradation by introducing two key innovations: (1) Token Refinement and Importance-guided Modeling (TRIM) module, which selects and refines informative tokens via SVD-based contribution scoring and cross-modal relevance modeling; (2) Per-sample Online Hard Negative Mining (PO-HNM) strategy dynamically maintains sample-specific negatives based on dual-consistency difficulty tracking, which is superior in long-text scenarios where key semantics are distributed in scattered positions. Extensive experiments on long-text image retrieval, short-text image retrieval, zero-shot classification, and text-to-image generation demonstrate OneLIP's robustness and versatility across diverse input lengths, offering a faithful solution for long-text representation learning of CLIP. Renjie Pan 0001, Jiayan Song, Hua Yang 0001 |
AAAI | 3 |
| 2026 | Physics-Environment Interaction Network for dense crowd behavior recognition
Yanshan Zhou, Renjie Pan 0001, Pingrui Lai, Hua Yang 0001 |
Pattern Recognit. | 5 |
| 2026 | Psychology-informed safety attributes recognition in dense crowds
Yanshan Zhou, Renjie Pan 0001, Cunyan Li, Hua Yang 0001 |
Pattern Recognit. Lett. | 5 |
| 2026 | MG-LLaVA: Toward Multi-Granularity Visual Instruction TuningabstractMulti-modal large language models (MLLMs) have made significant strides in various visual understanding tasks. However, the majority of these models are constrained to process low-resolution images, which limits their effectiveness in perception tasks that necessitate detailed visual information. In our study, we present MG-LLaVA, an innovative MLLM that enhances the model’s visual processing capabilities by incorporating a multi-granularity vision flow, which includes low-resolution, high-resolution, and object-centric features. We propose the integration of an additional high-resolution visual encoder to capture fine-grained details, which are then fused with base visual features through a Conv-Gate fusion network. To further refine the model’s object recognition abilities, we incorporate object-level features derived from bounding boxes identified by offline detectors. Being trained solely on publicly available multimodal data through instruction tuning, MG-LLaVA demonstrates exceptional perception skills. We instantiate MG-LLaVA with a wide variety of language encoders, ranging from 3.8B to 34B, to evaluate the model’s performance comprehensively. Extensive evaluations across multiple benchmarks demonstrate that MG-LLaVA outperforms existing MLLMs of comparable parameter sizes, showcasing its remarkable efficacy. Xiangtai Li, Haodong Duan, Haian Huang, Kai Chen 0026, Hua Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | Match Any KeypointsabstractPrevious research on sparse feature matching typically involves a staged optimization process of keypoint detection, description, and matching. While it allows the network to adapt to specific inputs, it may limit the network's expressive capability and the overall architectural flexibility. In this study, we rethink the matching framework and propose to directly match any given keypoints, optimizing the matching network in an approximately end-to-end manner. To achieve this, firstly, we dynamically sample random positions within the images as assumed keypoints during training, allowing the network to explore a broader matching space. Secondly, we replace specific descriptors with high-efficiency sparse embeddings at multi levels of the image, facilitating the direct learning of underlying textures. Thirdly, we propose a novel and promising architecture, called Proposal-Guided TRansformer (PGTR), which aggregates context information from neighboring match proposals instead of searching globally with local features. PGTR works especially well under our training approach, and attain a synergistic advantage in terms of performance and efficiency. The overall pipeline achieves outstanding performance on various keypoints without any retraining, and can be flexibly reused when new keypoints emerge, making it valuable for real-world applications. Code will be available. Renjie Pan 0001, Jun Zhou 0007, Hua Yang 0001, Cunyan Li |
IEEE Trans. Image Process. | 4 |
| 2026 | Signed Relation Graph Based Dynamical Interacting System Modeling for Multi-Agent Trajectory PredictionabstractMany complex systems prevalent in nature and society, from particle physics systems to social networks and team sports, can be viewed as dynamical interacting systems. Understanding the underlying interactions of agents in the system is the key task for predicting future behaviors of agents, which can be applied in various applications, e.g., autonomous vehicles and smart video surveillance. Since the interaction patterns between agents in the system can be dynamic and heterogeneous rather than fixed and homogeneous, it is very challenging to model interacting systems. In this paper, we design a novel graph structure called Signed Relation Graph (SRG) to model dynamical interacting systems. Since collective behaviors are very common in real-world scenes, our method is a group based model that takes heterogeneous relationships between agents into consideration, and achieves jointly modeling inter-group interactions and intra-group interactions. To assign signs on SRG, an unsupervised method called Relationship Reasoning Network is proposed. The relationship categories are reasoned explicitly, which makes handling multi-agent systems with multiple and dynamic interactions available. Further, Group Interaction Attention Graph Neural Network is proposed to aggregate information on SRG, which achieves not only reasoning the intensity of different interaction patterns but also modeling the trade-off between inter-group interactions and intra-group interactions. Our interacting systems modeling method can be used to predict multi-agent future trajectories in a variety of scenes with hard scenarios, including dense and drastic scenarios. Experimental results on three widely used human trajectory prediction datasets, including ETH and UCY in traffic scenes and NBA SportVU in sports scenes, demonstrate the effectiveness of our proposed model. Cunyan Li, Hua Yang 0001, Jun Sun 0005 |
IEEE Trans. Multim. | 2 |
| 2025 | OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human PreferenceabstractXiangyu Zhao, Shengyuan Ding, Zicheng Zhang, Haian Huang, Maosongcao Maosongcao, Jiaqi Wang, Weiyun Wang, Xinyu Fang, Wenhai Wang, Guangtao Zhai, Hua Yang, Haodong Duan, Kai Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Shengyuan Ding, Haian Huang, Maosongcao, Jiaqi Wang 0003, Weiyun Wang, Xinyu Fang, Wenhai Wang, Guangtao Zhai, Hua Yang 0001, Haodong Duan, Kai Chen 0026 |
ACL (1) | 11 |
| 2025 | LoKi: Low-dimensional KAN for Efficient Fine-tuning Image Modelsabstract’Pre-training + fine-tuning’ has been widely used in various downstream tasks. Parameter-efficient fine-tuning (PEFT) has demonstrated higher efficiency and promising performance compared to traditional full-tuning. The widely used adapter-based and prompt-based methods in PEFT can be uniformly represented as adding an MLP structure to the pre-trained model. These methods are prone to over-fitting in downstream tasks, due to the difference in data scale and distribution. To address this issue, we propose a new adapter-based PEFT module, i.e., LoKi, which consists of an encoder, a learnable activation layer, and a decoder. To maintain the simplicity of LoKi, we use single-layer linear networks for the encoder and decoder, and for the learnable activation layer, we use a Kolmogorov-Arnold Network (KAN) with the minimal number of layers (only 2 KAN linear layers). With a bottleneck rate much lower than that of Adapter, LoKi is equipped with fewer parameters (only half of Adapter) and eliminates slow training speed and high memory usage of KAN. We conduct extensive experiments on LoKi under image classification and video action recognition across 9 datasets. LoKi demonstrates highly competitive generalization performance compared to other PEFT methods with fewer tunable parameters, ensuring both effectiveness and efficiency. Xuan Cai, Renjie Pan 0001, Hua Yang 0001 |
CVPR | 3 |
| 2025 | Discovering Clone Negatives via Adaptive Contrastive Learning for Image-Text MatchingabstractIn this paper, we identify a common yet challenging issue in image-text matching, i.e., clone negatives: negative image-text pairs that semantically resemble positive pairs, leading to ambiguous and sub-optimal matching outcomes. To tackle this issue, we propose Adaptive Contrastive Learning (AdaCL), which introduces two margin parameters along with a modulating anchor to dynamically strengthen the compactness between positives and mitigate the influence of clone negatives. The modulating anchor is selected based on the distribution of negative samples without the need for explicit training, allowing for progressive tuning and advanced in-batch supervision. Extensive experiments across several tasks demonstrate the effectiveness of AdaCL in image-text matching. Furthermore, we extend AdaCL to weakly-supervised image-text matching by replacing human-annotated descriptions with automatically generated captions, thereby increasing the number of potential clone negatives. AdaCL maintains robustness in this setting, alleviating the reliance on crowd-sourced annotations and laying a foundation for scalable vision-language contrastive learning. Renjie Pan 0001, Jihao Dong, Hua Yang 0001 |
ICLR | 3 |
| 2025 | Cognitive Inspired Generalization Boosting for Face Forgery DetectionabstractCurrent face forgery detection models struggle to generalize across diverse forgery types due to limitations in feature extraction and refinement. Drawing inspiration from cognitive psychology, we propose an end-to-end detection framework SFMM to boost generalization capabilities. Based on dual-process theory, human cognition involves two complementary learning processes. The Fast Learning process rapidly captures explicit forgery-specific features using the Spatial-Frequency Transformer. The Slow Learning process gradually formulates generalized patterns across various forgeries by continuously updating a Momentum Memory Bank with the extracted features. This generalized knowledge is then employed to refine the feature extraction process. The two reciprocal processes are jointly optimized to foster a comprehensive and robust face forgery detection capability. Extensive experiments on benchmark datasets demonstrate that SFMM surpasses state-of-the-art methods in generalization tasks, achieving 1.0% and 3.7% AUC improvement on GAN-based and diffusion-based forgery datasets respectively. Yunwen Huang, Hua Yang 0001 |
ICME | 2 |
| 2025 | HAVEN: From Human Guidance to Assistant by Evolution Network in Vision-and-Language NavigationabstractVision-and-Language Navigation (VLN) is a critical task that enables robots to comprehend human instructions. The premise of current VLN tasks is built on the human’s familiarity with the environment structure, guiding the agent to complete the navigation task with explicit instructions, referred to as Human Guidance VLN (HG-VLN). However, real-world scenarios often involve humans unfamiliar with new environments, relying on agents to assist with navigation. In such tasks, humans can only provide destination-related information, requiring the agent to perform the path planning. We term this scenario Human Assistant VLN (HA-VLN). HA-VLN poses greater demands on the agent, therefore, we have restructured the classic Room to Room (R2R) dataset to introduce the Room to Room Assistant (R2RA) dataset, tailored for HA-VLN tasks. To address the challenges existing methods face when processing HA-VLN task instructions, we propose HAVEN: Human Assistant Vision-and- Language Navigation Evolution Network. This network integrates a Large Language Model (LLM) with an embedded memory system, achieving the paradigm shift from mainstream VLN methods to the HA-VLN task without requiring additional information. Our experiments demonstrate that algorithms incorporating HAVEN can achieve higher success rates in reaching destinations, shorter path selection, and lower navigation error rates in HA-VLN tasks. HAVEN can be integrated as an end-to-end module into any VLN method. The code and dataset is available at: https://github.com/longziyu/R2RA-Dataset Pingrui Lai, Zihao Xie, Hua Yang 0001 |
IJCNN | 3 |
| 2025 | Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual EditingabstractLarge Multi-modality Models (LMMs) have made significant progress in visual understanding and generation, but they still face challenges in General Visual Editing, particularly in following complex instructions, preserving appearance consistency, and supporting flexible input formats. To study this gap, we introduce RISEBench, the first benchmark for evaluating Reasoning-Informed viSual Editing (RISE). RISEBench focuses on four key reasoning categories: Temporal, Causal, Spatial, and Logical Reasoning. We curate high-quality test cases for each category and propose an robust evaluation framework that assesses Instruction Reasoning, Appearance Consistency, and Visual Plausibility with both human judges and the LMM-as-a-judge approach. We conducted experiments evaluating nine prominent visual editing models, comprising both open-source and proprietary models. The evaluation results demonstrate that current models face significant challenges in reasoning-based editing tasks. Even the most powerful model evaluated, GPT-image-1, achieves an accuracy of merely 28.8%. RISEBench effectively highlights the limitations of contemporary editing models, provides valuable insights, and indicates potential future directions for the field of reasoning-aware visual editing. Our code and data have been released at https://github.com/PhoenixZ810/RISEBench. Peiyuan Zhang, Kexian Tang, Xiaorong Zhu, Hao Li 0069, Wenhao Chai, Renqiu Xia, Guangtao Zhai, Junchi Yan, Hua Yang 0001, Xue Yang 0005, Haodong Duan |
NeurIPS | 11 |
| 2025 | ReAL: Improving Image-Text Retrieval with Authentic Negative Repository LearningabstractCurrent methods for image-text retrieval commonly propose various fusion modules to achieve robust visual-textual alignment, primarily relying on in-batch learning to guide the matching process. Some follow-up methods seek to enlarge the number of negative samples to boost image-text contrastive learning. However, these methods often face challenges posed by semantic-consistent negatives, i.e., negative samples that share correspondence with the ground truth, leading to confusion in learning cross-modal semantics. To address this issue, we propose a novel Retrieve with Authentic Negative Repository Learning (ReAL) method, which constructs a specific Authentic Negative Repository filled with valuable negative sample pairs. By introducing a Unique Negative Filter with a Discriminative Triplet Ranking Loss, ReAL effectively filters out the semantic-consistent negatives through similarity distribution analysis and threshold learning. Moreover, existing fusion paradigms suffer from intricate use of fine-grained representations from word- and region-level instances to progressively refine the fused embedding. In this article, we propose a lightweight Cluster Refinement Module to exploit cross-modal semantics in a 1-way-1-out paradigm. Each visual-textual alignment can spontaneously uncover correlations with adjacent alignments through aggregation and re-allocation, without the need for a redundant and cost-inefficient refinement stage. Furthermore, ReAL employs dual momentum encoders with two memory banks, expanding the selection range of the Authentic Negative Repository to include a broader set of negatives. Extensive experiments conducted on Flickr30K, MS-COCO, and the augmented Flickr30K (with more hard negatives) demonstrate the superiority and robustness of ReAL, while also showcasing its significantly reduced inference time compared to other competitive baselines. Renjie Pan 0001, Hua Yang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | M-RAT: a Multi-grained Retrieval Augmentation Transformer for Image Captioning
Jiayan Song, Renjie Pan 0001, Jun Zhou 0007, Hua Yang 0001 |
ACCV (3) | 4 |
| 2024 | FC-GNN: Recovering Reliable and Accurate Correspondences from InterferencesabstractFinding correspondences between images is essential for many computer vision tasks and sparse matching pipelines have been popular for decades. However, matching noise within and between images, along with inconsistent key-point detection, frequently degrades the matching performance. We review these problems and thus propose: 1) a novel and unified Filtering and Calibrating (FC) approach that jointly rejects outliers and optimizes inliers, and 2) leveraging both the matching context and the underlying image texture to remove matching uncertainties. Under the guidance of the above innovations, we construct Filtering and Calibrating Graph Neural Network (FC-GNN), which follows the FC approach to recover reliable and accurate correspondences from various interferences. FC-GNN conducts an effectively combined inference of contextual and local information through careful embedding and multiple information aggregations, predicting confidence scores and calibration offsets for the input correspondences to jointly filter out outliers and improve pixel-level matching accuracy. Moreover, we exploit the local coherence of matches to perform inference on local graphs, thereby reducing computational complexity. Overall, FC-GNN operates at lightning speed and can greatly boost the performance of diverse matching pipelines across various tasks, showcasing the immense potential of such approaches to become standard and pivotal components of image matching. Code is avaiable at https://github.com/xuy123456/fcgnn. Jun Zhou 0007, Hua Yang 0001, Renjie Pan 0001, Cunyan Li |
CVPR | 3 |
| 2024 | Exploring Interactive Semantic Alignment for Efficient HOI Detection with Vision-language ModelabstractHuman-Object Interaction (HOI) detection aims to localize human-object pairs and comprehend their interactions. Recently, two-stage transformer-based methods have demonstrated competitive performance. However, these methods frequently focus on object appearance features and ignore global contextual information. Besides, vision-language model CLIP which effectively aligns visual and text embeddings has shown great potential in zero-shot HOI detection. Based on the former facts, We introduce a novel HOI detector named ISA-HOI, which extensively leverages knowledge from CLIP, aligning interactive semantics between visual and textual features. We first extract global context of image and local features of object to Improve interaction Features in images (IF). On the other hand, we propose a Verb Semantic Improvement (VSI) module to enhance textual features of verb labels via cross-modal fusion. Ultimately, our method achieves competitive results on the HICO-DET and V-COCO benchmarks with much fewer training epochs, and outperforms the state-of-the-art under zero-shot settings. Jihao Dong, Hua Yang 0001, Renjie Pan 0001 |
ICME | 2 |
| 2024 | Hydrodynamics-Informed Neural Network for Simulating Dense Crowd Motion PatternsabstractWith global occurrences of crowd crushes and stampedes, dense crowd simulation has been drawing great attention. In this research, our goal is to simulate dense crowd motions under six classic motion patterns, more specifically, to generate subsequent motions of dense crowds from the given initial states. Since dense crowds share similarities with fluids, such as continuity and fluidity, one common approach for dense crowd simulation is to construct hydrodynamics-based models, which consider dense crowds as fluids, guide crowd motions with Navier-Stokes equations, and conduct dense crowd simulation by solving governing equations. Despite the proposal of these models, dense crowd simulation faces multiple challenges, including the difficulty of directly solving Navier-Stokes equations due to their nonlinear nature, the ignorance of distinctive crowd characteristics which fluids lack, and the gaps in the evaluation and validation of crowd simulation models. To address the above challenges, we build a hydrodynamic model, which captures the crowd physical properties (continuity, fluidity, etc.) with Navier-Stokes equations and reflects the crowd social properties (sociality, personality, etc.) with operators that describe crowd interactions and crowd-environment interactions. To tackle the computational problem, we propose to solve the governing equation based on Navier-Stokes equations using neural networks, and introduce the Hydrodynamics-Informed Neural Network (HINN) which preserves the structure of the governing equation in its network architecture. To facilitate the evaluation, we construct a new dense crowd motion video dataset called Dense Crowd Flow Dataset (DCFD), containing six classic motion patterns (line, curve, circle, cross, cluster and scatter) and 457 video clips, which can serve as the groundtruths for various objective metrics. Numerous experiments are conducted using HINN to simulate dense crowd motions under six motion patterns with video clips from DCFD. Objective evaluation metrics that concerns authenticity, fidelity and diversity demonstrate the superior performance of our model in dense crowd simulation compared to other simulation models. Our code and dataset are available at https://github.com/shanshan-zys/HINN. Yanshan Zhou, Pingrui Lai, Yingjie Xiong, Hua Yang 0001 |
ACM Multimedia | 5 |
| 2024 | Joint Intra & Inter-Grained Reasoning: A New Look Into Semantic Consistency of Image-Text RetrievalabstractMultimodal understanding aims at constructing semantic correlations among modalities of data while performing various downstream tasks. As one of the primary multimodal downstream tasks, image-text retrieval imposes a high demand on semantic alignment because of the independent expression paradigms of images and text. Existing methods mainly construct a joint embedding space at a single granularity level (either global or local). However, such single reasoning paradigms lack granularity interaction, resulting in semantic inconsistency and cross-domain catastrophes. To address these issues, we design a novel Joint Intra and Inter-grained Network (JIIGNet), focusing on not only intra- but also inter-grained interaction between modalities by combining scene information (global) with region-level (local) instances. Specifically, we simultaneously initiate three specific alignment modules, i.e., global-grained, local-grained, and cross-grained alignment modules, followed by Triplet Attention Refinement to better refine the fused embedding at the alignment-level with proper self and cross attention. For different scenarios, a Style Adaptation Head is further designed to smartly accommodate different samples. We validate JIIGNet through extensive experiments conducted on two widely used datasets: Flickr-30 K and MS-COCO, demonstrating the effectiveness of our proposed method. Renjie Pan 0001, Hua Yang 0001, Cunyan Li, Jinhai Yang 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Psychology-Guided Environment Aware Network for Discovering Social Interaction Groups from VideosabstractSocial interaction is a common phenomenon in human societies. Different from discovering groups based on the similarity of individuals’ actions, social interaction focuses more on the mutual influence between people. Although people can easily judge whether or not there are social interactions in a real-world scene, it is difficult for an intelligent system to discover social interactions. Initiating and concluding social interactions are greatly influenced by an individual’s social cognition and the surrounding environment, which are closely related to psychology. Thus, converting the psychological factors that impact social interactions into quantifiable visual representations and creating a model for interaction relationships poses a significant challenge. To this end, we propose a Psychology-Guided Environment Aware Network (PEAN) that models social interaction among people in videos using supervised learning. Specifically, we divide the surrounding environment into scene-aware visual-based and human-aware visual-based descriptions. For the scene-aware visual clue, we utilize 3D features as global visual representations. For the human-aware visual clue, we consider instance-based location and behaviour-related visual representations to map human-centred interaction elements in social psychology: distance, openness, and orientation. In addition, we design an environment aware mechanism to integrate features from visual clues, with a Transformer to explore the relation between individuals and construct pairwise interaction strength features. The interaction intensity matrix reflecting the mutual nature of the interaction is obtained by processing the interaction strength features with the interaction discovery module. An interaction constrained loss function composed of interaction critical loss function and smoothFβloss function is proposed to optimize the whole framework to improve the distinction of the interaction matrix and alleviate class imbalance caused by pairwise interaction sparsity. Given the diversity of real-world interactions, we collect a new dataset named Social Basketball Activity Dataset (Soical-BAD), covering complex social interactions. Our method achieves the best performance among social-CAD, social-BAD, and their combined dataset named Video Social Interaction Dataset (VSID). Jinhai Yang 0001, Hua Yang 0001, Renjie Pan 0001, Pingrui Lai, Guangtao Zhai |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | Sine: Similarity-Regularized Intra-Class Exploitation for Cross-Granularity Few-Shot LearningabstractFew-shot learning aims for rapid adaptation with few samples. Recently, cross-granularity few-shot learning has emerged as a promising research area, where models observe coarse labels but target fine-grained recognition among novel classes. As coarse supervision tends to eliminate feature discrimination among underlying sub-classes, existing methods commonly utilize self-supervision as a complement to explore intra-class variation. However, current methods suffer from an intrinsic conflict between contrastive learning and coarse supervision. In this paper, we locate the root cause of the intrinsic conflict. Then, we resolve it by exploiting the similarity among augmented views while ignoring the unreasonable constraint between negative pairs. Besides, we decouple contrastive learning and coarse supervision into parallel branches to better regularize the latent space. Albeit simple, our approach consistently outperforms state-of-the-art methods across different benchmarks. Jinhai Yang 0001, Hua Yang 0001 |
ICASSP | 2 |
| 2023 | Fine-Grained Music Plagiarism Detection: Revealing Plagiarists through Bipartite Graph Matching and a Comprehensive Large-Scale DatasetabstractMusic plagiarism detection is gaining more and more attention due to the popularity of music production and society's emphasis on intellectual property. We aim to find fine-grained plagiarism in music pairs since conventional methods are coarse-grained and cannot match real-life scenarios. Considering that there is no sizeable dataset designed for the music plagiarism task, we establish a large-scale simulated dataset, named Music Plagiarism Detection Dataset (MPD-Set) under the guidance and expertise of researchers from national-level professional institutions in the field of music. MPD-Set considers diverse music plagiarism cases found in real life from the melodic, rhythmic, and tonal levels respectively. Further, we establish a Real-life Dataset for evaluation, where all plagiarism pairs are real cases. To detect the fine-grained plagiarism pairs effectively, we propose a graph-based method called Bipatite Melody Matching Detector (BMM-Det), which formulates the problem as a max matching problem in the bipartite graph. Experimental results on both the simulated and Real-life Datasets demonstrate that BMM-Det outperforms the existing plagiarism detection methods, and is robust to common plagiarism cases like transpositions, pitch shifts, duration variance, and melody change. Datasets and source code are open-sourced at https://github.com/xuan301/BMMDet_MPDSet. Wenxuan Liu 0004, Tianyao He, Chen Gong 0006, Ning Zhang 0040, Hua Yang 0001, Junchi Yan |
ACM Multimedia | 5 |
| 2023 | Matching-to-Detecting: Establishing Dense and Reliable Correspondences Between Images
Jun Zhou 0007, Renjie Pan 0001, Hua Yang 0001, Cunyan Li |
PRCV (2) | 4 |
| 2023 | Adaptive and Collaborative Multi-scale Alignment for Text-Based Person SearchabstractText-to-image person search is challenging due to the cross-scale correspondences and information inequality between modalities. Specifically, images and text are complexly linked at different scales and images are usually more informative and complete than text. It is crucial to establish semantic correlations between modalities and focus on task-relevant information in images. In this paper, we propose a novel Adaptive and Collaborative Multi-scale Alignment network (ACMA) for text-based person search that learns semantically consistent and information-aligned multi-modal representations. Firstly, we introduce a novel joint embedding module that adaptively integrates features of different pixels and words, thereby extracting semantically consistent multi-modal features at different scales. Second, we design a cross-modal fusion feature-based auxiliary visual branch to guide the extraction of key visual features that are beneficial for cross-modal matching. Extensive experiments validate that ACMA outperforms the state-of-the-art method. Renjie Pan 0001, Hua Yang 0001 |
VCIP | 3 |
| 2022 | Subtask-dominated Supervised Pretraining Transfer Learning for Person Search
Hua Yang 0001, Shibao Zheng |
BMVC | 2 |
| 2022 | Mask-Guided Transformer for Human-Object Interaction DetectionabstractHuman-object interaction (HOI) detection is a meaningful research topic on human activity understanding. Recent works have made significant progress by focusing on efficient triplet matching and leveraging image-wide features based on encoder-decoder architecture. However, the ability to gather relevant contextual information about human is limited and different sub-tasks in HOI detection are not differentiated by specific decoupling in previous methods. To this end, we propose a new transformer-based method for HOI detection, namely, Mask-Guided Transformer (MGT). Our model, which is composed of five parallel decoders with a shared encoder, not only emphasizes interactive regions by applying body features, but also disentangles the prediction of instance and interaction. We achieve a favorable result at 63.3 mAP on the well-known HOI detection dataset V-COCO. Daocheng Ying, Hua Yang 0001, Jun Sun 0005 |
VCIP | 2 |
| 2022 | Modeling context appearance changes for person re-identification via IPES-GCN
Hua Yang 0001, Ji Zhu 0002, Qin Zhou 0002, Shibao Zheng |
Neurocomputing | 2 |
| 2022 | Making person search enjoy the merits of person re-identification
Hua Yang 0001, Qin Zhou 0002, Shibao Zheng |
Pattern Recognit. | 2 |
| 2021 | PRN: Psychology-Inspired Relation Network for Detecting Social Interaction Groups from Single Images
Jinhai Yang 0001, Hua Yang 0001, Guangtao Zhai |
BMVC | 3 |
| 2021 | Attentive Semantic Exploring for Manipulated Face DetectionabstractFace manipulation methods develop rapidly in recent years, whose potential risk to society accounts for the emerging of researches on detection methods. However, due to the diversity of manipulation methods and the high quality of fake images, detection methods suffer from a lack of generalization ability. To solve the problem, we find that segmenting images into semantic fragments could be effective, as discriminative defects and distortions are closely related to such fragments. Besides, to highlight discriminative regions in fragments and to measure contribution to the final prediction of each fragment is efficient for the improvement of generalization ability. Therefore, we propose a novel manipulated face detection method based on Multilevel Facial Semantic Segmentation and Cascade Attention Mechanism. To evaluate our method, we reconstruct two datasets: GGFI and FFMI, and also collect two open-source datasets. Experiments on four datasets verify the advantages of our approach against other state-of-the-arts, especially its generalization ability. Hua Yang 0001 |
ICASSP | 2 |
| 2021 | Crowd Counting Via Multi-Level Regression With Latent Gaussian MapsabstractCrowd counting still confronts two primary challenges: limited ability to deal with cross density levels caused by fixed density maps and lack of fine-grained or coarse-grained guidance for density estimation. In this paper, a novel end-to-end crowd counting framework via multi-level regression with latent Gaussian maps is proposed, which is consisted of GaussianNet, EstimateNet and Discriminator. GaussianNet is composed of masked Gaussian convolutional blocks and vanillia convolutional layers, to generate latent Gaussian maps adaptively for various density levels. The latent Gaussian maps are then treated as the ground truth density maps for EstimateNet, which outputs density estimations and follows the principle of adversarial learning with Discriminator. Moreover, multi-level losses are combined for density map regression guidance. Extensive experiments on the major public datasets outperform state-of-the-art ones, illustrating the superior validity of the proposed framework. Yukang Gao, Hua Yang 0001 |
ICASSP | 2 |
| 2021 | MPASNET: Motion Prior-Aware Siamese Network For Unsupervised Deep Crowd Segmentation In Video ScenesabstractCrowd segmentation is a fundamental task serving as the basis of crowded scene analysis, and it is highly desirable to obtain refined pixel-level segmentation maps. However, it remains a challenging problem, as existing approaches either require dense pixel-level annotations to train deep learning models or merely produce rough segmentation maps from optical or particle flows with physical models. In this paper, we propose the Motion Prior-Aware Siamese Network (MPASNET) for unsupervised crowd semantic segmentation. This model not only eliminates the need for annotation but also yields high-quality segmentation maps. Specially, we first analyze the coherent motion patterns across the frames and then apply a circular region merging strategy on the collective particles to generate pseudo-labels. Moreover, we equip MPASNET with siamese branches for augmentation-invariant regularization and siamese feature aggregation. Experiments over benchmark datasets indicate that our model outperforms the state-of-the-arts by more than 12% in terms of mIoU. Jinhai Yang 0001, Hua Yang 0001 |
ICIP | 2 |
| 2021 | Adversarial Cross-Scale Alignment Pursuit For Seriously Misaligned Person Re-IdentificationabstractPerson re-identification is still a challenging task in actual application due to serious misalignment caused by large scale variation, serious occlusion, or variably truncated body. Conventional holistic methods usually lack the cross-scale aligning ability. Segmentation-based partial methods achieve better aligning performance but generally suffer from the instability of part segmentation. To explicitly address these issues, we define the seriously misaligned Re-ID task and propose a novel framework called adversarial cross-scale alignment pursuit (ACSAP). Instead of dynamically segmenting the feature map for part alignment, the model incorporates the stability of holistic methods and adversarially generates aligned feature maps for similarity metrics. Especially, the spatial reconstruction (SR) module in the generator is proposed for feature filtering and aligning. Then a part visibility calculation (PVC) algorithm is proposed to distinguish the credibility of different generated areas. We propose a novel dataset called Seriously-Misaligned-REID and achieve 10.0% rank-1 outperformance compared to state-of-the-art methods on it. Extensive performance results demonstrate the effectiveness of our framework. Yuanhang He, Hua Yang 0001, Lin Chen 0019 |
ICIP | 2 |
| 2021 | Towards Cross-Granularity Few-Shot Learning: Coarse-to-Fine Pseudo-Labeling with Visual-Semantic Meta-EmbeddingabstractFew-shot learning aims at rapidly adapting to novel categories with only a handful of samples at test time, which has been predominantly tackled with the idea of meta-learning. However, meta-learning approaches essentially learn across a variety of few-shot tasks and thus still require large-scale training data with fine-grained supervision to derive a generalized model, thereby involving prohibitive annotation cost. In this paper, we advance the few-shot classification paradigm towards a more challenging scenario, i.e, cross-granularity few-shot classification, where the model observes only coarse labels during training while is expected to perform fine-grained classification during testing. This task largely relieves the annotation cost since fine-grained labeling usually requires strong domain-specific expertise. To bridge the cross-granularity gap, we approximate the fine-grained data distribution by greedy clustering of each coarse-class into pseudo-fine-classes according to the similarity of image embeddings. We then propose a meta-embedder that jointly optimizes the visual- and semantic-discrimination, in both instance-wise and coarse class-wise, to obtain a good feature space for this coarse-to-fine pseudo-labeling process. Extensive experiments and ablation studies are conducted to demonstrate the effectiveness and robustness of our approach on three representative datasets. Jinhai Yang 0001, Hua Yang 0001, Lin Chen 0019 |
ACM Multimedia | 2 |
| 2021 | Complex Event Recognition via Spatial-Temporal Relation Graph ReasoningabstractEvents in videos usually contain a variety of factors: objects, environments, actions, and their interaction relations, and these factors as the mid-level semantics can bridge the gap between the event categories and the video clips. In this paper, we present a novel video events recognition method that uses the graph convolution networks to represent and reason the logic relations among the inner factors. Considering that different kinds of events may focus on different factors, we especially use the transformer networks to extract the spatial-temporal features drawing upon the attention mechanism that can adaptively assign weights to concerned key factors. Although transformers generally rely more on large datasets, we show the effectiveness of applying a 2D convolution backbone before the transformers. We train and test our framework on the challenging video event recognition dataset UCF-Crime and conduct ablation studies. The experimental results show that our method achieves state-of-the-art performance, outperforming previous principal advanced models with a significant margin of recognition accuracy. Hongtian Zhao, Hua Yang 0001 |
VCIP | 3 |
| 2021 | Harmonious attention network for person re-identification via complementarity between groups and individuals
Lin Chen 0019, Hua Yang 0001, Qiling Xu |
Neurocomputing | 2 |
| 2021 | Graph similarity rectification for person search
Hua Yang 0001, Ji Zhu 0002, Xinzhe Li 0002, Zhigang Chang, Shibao Zheng |
Neurocomputing | 2 |
| 2021 | Robust and Efficient Graph Correspondence Transfer for Person Re-IdentificationabstractSpatial misalignment caused by variations in poses and viewpoints is one of the most critical issues that hinder the performance improvement in existing person re-identification (Re-ID) algorithms. Although it is straightforward to explore correspondence learning algorithms for alignment, online learning is intractable for negative pairs due to the intrinsic visual difference between negative pairs and efficiency concern. To address this problem, in this paper, we present a robust and efficient graph correspondence transfer (REGCT) approach for explicit spatial alignment in Re-ID. Specifically, we propose the off-line correspondence learning and on-line correspondence transfer framework. During training, patch-wise correspondences between positive training pairs are established via graph matching. By exploiting both spatial and visual contexts of human appearance in graph matching, meaningful semantic correspondences can be obtained. During testing, the off-line learned patch-wise correspondence templates are transferred to test pairs with similar pose-pair configurations for local feature distance calculation. To enhance the robustness of correspondence transfer, we design a novel pose context descriptor to accurately model human body configurations, and present an approach to measure the similarity between a pair of pose context descriptors. Meanwhile, to improve testing efficiency, we propose a correspondence template ensemble method using the voting mechanism, which significantly reduces the amount of patch-wise matchings involved in distance calculation. With the aforementioned strategies, the REGCT model can effectively and efficiently handle the spatial misalignment problem in Re-ID. Extensive experiments on five challenging benchmarks, including VIPeR, Road, PRID450S, 3DPES, and CUHK01, evidence the superior performance of REGCT over other state-of-the-art approaches. Qin Zhou 0002, Heng Fan 0001, Hua Yang 0001, Hang Su 0006, Shibao Zheng, Shuang Wu 0001, Haibin Ling |
IEEE Trans. Image Process. | 3 |
| 2021 | Group Re-Identification With Group Context Graph Neural NetworksabstractGroup re-identification aims to match groups of people across disjoint cameras. In this task, the contextual information from neighbor individuals can be exploited for re-identifying each individual within the group as well as the entire group. However, compared with single person re-identification, it brings new challenges including group layout and group membership changes. Motivated by the observation that individuals who are close together are more likely to keep in the same group under different cameras than those who are far apart, we propose to model each group as a spatial K-nearest neighbor graph (SKNNG) and design a group context graph neural network (GCGNN) for graph representation learning. Specifically, for each node in the graph, the proposed GCGNN learns an embedding which aggregates the contextual information from neighbor nodes. We design multiple weighting kernels for neighborhood aggregation based on the graph properties including node in-degrees and spatial relationship attributes. We compute the similarity scores between node embeddings of two graphs for group member association and obtain the matching score between the two graphs by summing up the similarity scores of all linked node pairs. Experimental results on three public datasets show that our approach performs favorably against state-of-the-art methods and achieves high efficiency. Ji Zhu 0002, Hua Yang 0001, Weiyao Lin, Nian Liu 0002, Jia Wang 0004, Wenjun Zhang 0001 |
IEEE Trans. Multim. | 2 |
| 2020 | Multilevel Interaction Reasoning For Complex Event RecognitionabstractEvent as a complex process contains many factors. Objects, environments, and their interactions vary with time. Recognizing event remains a challenging task in computer vision. In this paper, a multilevel interaction reasoning framework is proposed for complex event recognition. Firstly, we construct a 3D ConvNet to extract the spatial-temporal feature to present global scene. Then a graph ConvNet is built to reasoning about the multilevel interaction: object-object and object-environment interaction, via graphs that contain object feature in video, and projection of global scene feature in 3D ConvNet. The proposed method effectively explore the nature of events occurring and developing, and experimental results on challenging UCF-Crime dataset achieve state-of-the-art with a 3.5% gain over other models. Hua Yang 0001, Jun Sun 0005 |
ICIP | 2 |
| 2020 | Automatic Region Selection For Objective Sharpness Assessment Of Mobile Device PhotosabstractMobile devices are the source of a vast majority of digital photos today. Photos taken by mobile devices generally have fairly good visual quality. When evaluating high-quality mobile device photos, people have to manually zoom in to local regions to discern the subtle difference. Understandably, a global objective quality assessment method cannot perform well on such task. Therefore, local region selection is widely recognized as a prerequisite for the following quality evaluation. Clearly, subjective regions selection suffers from the drawbacks in terms of productivity, reproducibility and optimality. In this paper, we propose an automatic local region selection algorithm for sharpness measurement of mobile device photos. Specifically, local texture statistics, depth, saliency, as well as inter-pictures difference, are used as main features to select an optimal local region, in which the sharpness is then measured. For validation, we have built a largescale database for sharpness evaluation of mobile device photos, with 100 different scenes shot by several flagship mobile phones. The experimental results show that the performance of classic sharpness evaluation algorithms can be substantially improved with the region selected by the proposed algorithm. Guangtao Zhai, Wenhan Zhu, Yucheng Zhu, Xiongkuo Min, Xiao-Ping Zhang 0002, Hua Yang 0001 |
ICIP | 7 |
| 2020 | Dual-Mode iterative denoiser: Tackling the weak label for anomaly detectionabstractCrowd anomaly detection suffers from limited training data under weak supervision. In this paper, we propose a dual-mode iterative denoiser to tackle the weak label challenge for anomaly detection. First, we use a convolution autoencoder (CAE) in image space to act as a cluster for grouping similar video clips, where the spatial-temporal similarity helps the cluster metric to represent the reconstruction error. Then we use the graph convolution neural network (GCN) to explore the temporal correlation and the feature similarity between video clips within different rough labels, where the classifier can be constantly updated in the label denoising process. Without specific image-level labels, our model can predict the clip-level anomaly probabilities for videos. Extensive experiment results on two public datasets show that our approach performs favorably against the state-of-the-art methods. Shuheng Lin, Hua Yang 0001 |
ICPR | 2 |
| 2020 | Weakly Supervised Pedestrian Attribute Recognition with Attention in Latent Space
Mingjun Sun, Hua Yang 0001, Guangtao Zhai |
PRCV (2) | 2 |
| 2020 | Weighted bilinear coding over salient body parts for person re-identification
Zhigang Chang, Heng Fan 0001, Hang Su 0006, Hua Yang 0001, Shibao Zheng, Haibin Ling |
Neurocomputing | 5 |
| 2020 | Person re-identification from virtuality to reality via modality invariant adversarial mechanism
Lin Chen 0019, Hua Yang 0001 |
Neurocomputing | 2 |
| 2020 | Comprehensive feature fusion mechanism for video-based person re-identification via significance-aware attention
Lin Chen 0019, Hua Yang 0001 |
Signal Process. Image Commun. | 2 |
| 2019 | Social MIL: Interaction-Aware for Crowd Anomaly DetectionabstractCrowd anomaly detection under surveillance scene is a quite challenging task, which often companies with not rare objects, unexpected bursts in activity and complex dynamic patterns. In this paper, we propose a social multiple-instance learning(MIL) framework with a dual-branch network by considering dynamic interaction among groups, individuals and environment to obtain attentive spatial-temporal feature representation. First, MIL is employed to overcome the challenge of rare training abnormal samples and video-based labels. The social force map is utilized for modeling behavior interaction to supply the prior knowledge. In addition, we introduce the self-attention module, which represents a more discriminative spatial-temporal feature based on C3D network through implementing weight redistribution inside the feature. The results of the experiments conducted on UCF-Crime dataset show that the proposed dual-branch social multiple-instance learning (MIL) anomaly detection framework with the dual-branch network outperforms than existing approaches and obtains the state-of-the-art performance. Shuheng Lin, Hua Yang 0001, Xianchao Tang, Tianqi Shi, Lin Chen 0019 |
AVSS | 2 |
| 2019 | An Online Crowd Semantic Segmentation Method Based on Reinforcement LearningabstractIn this paper, we propose an online crowd segmentation method based on reinforcement learning. Existing approaches for the task suffers from the fixed parameters applied to different sceneries. We propose to utilize the reinforcement learning to adaptively adjust the critical parameters for better clustering. Specifically, we first propose the velocity-constrained natural nearest neighbor algorithm to adaptively determine the original fixed parameter K in K-nearest-neighbor, thus better reflecting the neighborhood relationship in velocity correlation calculation. Then, we construct an online reinforcement learning module in the grouping process to adaptively decide the segmentation threshold. Furthermore, we combine the motion image semantics with feature point grouping to obtain the pixel-level segmentation. This method enhances the generalization and robustness in different scenarios to increase the accuracy. Comprehensive experimental results conducted on different datasets demonstrate the effectiveness and validity of our method. Hua Yang 0001, Lin Chen 0019 |
ICIP | 2 |
| 2019 | Group Re-Identification with Hybrid Attention Model and Residual DistanceabstractGroup re-identification (Re-ID) is an important task of computer vision and involves multiple challenges. In this paper, we propose a Hybrid Attention Model (HAM) to address the problem of group Re-ID, solving the spatial variation in the challenging still-image-based group Re-ID task. HAM consists of both the position and channel attention to make the network focus more on the crucial areas and features of the group images. Furthermore, we propose a novel Least Squares Residual Distance (LSRD) based on the least squares algorithm. LSRD can leverage the residual of the fitting function achieved by least squares method, better learning the metric between group image pairs. To evaluate the performance, we propose a new largest group Re-ID dataset. Extensive experimental results conducted on the dataset demonstrate the effectiveness of our approaches. Qiling Xu, Hua Yang 0001, Lin Chen 0019, Guangtao Zhai |
ICIP | 2 |
| 2019 | Scalable Receptive Field GAN: An End-to-End Adversarial Learning Framework for Crowd Counting
Yukang Gao, Hua Yang 0001 |
PRCV (2) | 2 |
| 2019 | Spatial-temporal Fusion Network with Residual Learning and Attention Mechanism: A Benchmark for Video-Based Group Re-ID
Qiling Xu, Hua Yang 0001, Lin Chen 0019 |
PRCV (1) | 2 |
| 2019 | Distribution Context Aware Loss for Person Re-identificationabstractTo learn the optimal similarity function between probe and gallery images in Person re-identification, effective deep metric learning methods have been extensively explored to obtain discriminative feature embedding. However, existing metric loss like triplet loss and its variants always emphasize pair-wise relations but ignore the distribution context in feature space, leading to inconsistency and sub-optimal. In fact, the similarity of one pair not only decides the match of this pair, but also has potential impacts on other sample pairs. In this paper, we propose a novel Distribution Context Aware (DCA) loss based on triplet loss to combine both numerical similarity and relation similarity in feature space for better clustering. Extensive experiments on three benchmarks including Market-1501, DukeMTMC-reID and MSMT17, evidence the favorable performance of our method against the corresponding baseline and other state-of-the-art methods. Zhigang Chang, Qin Zhou 0002, Shibao Zheng, Hua Yang 0001, Tai-Pang Wu |
VCIP | 5 |
| 2019 | Reranking optimization for person re-identification under temporal-spatial information and common network consistency constraints
Hua Yang 0001, Zhaoxi Cheng, Lin Chen 0019 |
Pattern Recognit. Lett. | 1 |
| 2018 | Online Multi-Object Tracking with Dual Matching Attention Networks
Ji Zhu 0002, Hua Yang 0001, Nian Liu 0002, Wenjun Zhang 0001, Ming-Hsuan Yang 0001 |
ECCV (5) | 2 |
| 2018 | Learning to Predict where the Children with Asd LookabstractAs is known to us, people with Autism Spectrum Disorder (ASD) have atypical visual attention towards stimuli. Learning the visual attention of people especially, children, with ASD contribute to related research in the field of medicine and psychology. In this paper, we first construct a saliency prediction for children with autism (SPCA) database, which is the first of its kind and consists of 500 images and the corresponding eye tracking data collected from 13 different children with ASD. We compare the performance of five state-of-the-art deep neural networks (DNN)-based saliency prediction approaches with their original networks and the fine-tuned networks on our database. We predict the atypical visual attention of children with ASD for the first time and get the best saliency prediction results for individuals with ASD so far. Huiyu Duan, Guangtao Zhai, Xiongkuo Min, Yi Fang 0009, Zhaohui Che, Xiaokang Yang 0001, Cheng Zhi, Hua Yang 0001 |
ICIP | 8 |
| 2018 | Spatial Invariant Person Search Network
Liangqi Li, Hua Yang 0001, Lin Chen 0019 |
PRCV (2) | 2 |
| 2018 | Comprehensive Samples Constrain for Person SearchabstractIn this paper, we propose a method to further improve person search by fully utilizing the combination of pedestrian detection and person re-identification tasks. An improved constrain that utilizes comprehensive samples in the dataset is proposed to fully excavate information for recognition. Besides the label constrain for training the model in traditional classification task, unlabeled identities that do not have specific IDs are utilized as well to constitute a tailored triplet loss for more performance improvement. Meanwhile, a novel large-scale challenging dataset, SJTU318, which uses videos acquired through twelve cameras is proposed to demonstrate the effectiveness of our method. It contains 443 identities and 14,610 frames in which pedestrians are annotated with their bounding box positions and identities. Experiments conducted on a public dataset, CUHK-SYSU and our proposed dataset SJTU318 show that our method outperforms existing state-of-the-art approaches. Liangqi Li, Hua Yang 0001, Lin Chen 0019 |
VCIP | 2 |
| 2017 | Data Generation for Improving Person Re-identificationabstractIn this paper, we explore ways to address the challenges such as data bias caused by the lack of data on person re-identification problem. We propose a data generation framework from both intra- and inter-view aspects for data augmentation to advance the performance of the existing person re-identification algorithms. Specifically, for intra-view data generation, the proposed method generates useful predicted sequences within a camera view for certain person data expansion. The generated sequences well preserve the movement information of the camera and objects, which expands the original data with longer sequence length to tackle the problem caused by insufficient data from the root. For more challenging datasets which suffer from background clutters, we propose an inter-view image generation with automatic end-to-end background substitution to eliminate the influence by the background and increase the diversity of the training data as well, which makes the recognition system learn to focus on the regions of objects and image features related to identity. We then propose a flexible data augmentation method based on our data generation approaches to improve the performance of the person re-identification and analyze the advantages and applicability of these approaches respectively. Evaluated on the challenging re-id datasets, our method outperforms existing state-of-the-art approaches without any network structure modification on the baseline neural network. Cross-datasets evaluation results show that our method has favorable generalization ability and is potentially helpful for solving similar recognition tasks due to the common issue of insufficient data. Lin Chen 0019, Hua Yang 0001, Shuang Wu 0001 |
ACM Multimedia | 2 |
| 2017 | Crowd Behavior Analysis via Curl and Divergence of Motion Trajectories
Shuang Wu 0001, Hua Yang 0001, Shibao Zheng, Hang Su 0006, Yawen Fan, Ming-Hsuan Yang 0001 |
Int. J. Comput. Vis. | 2 |
| 2017 | Edge guided salient object detection
Bing Yang 0003, Xiaoyun Zhang 0001, Li Chen 0021, Hua Yang 0001 |
Neurocomputing | 4 |
| 2017 | Bilinear dynamics for crowd video analysis
Shuang Wu 0001, Hang Su 0006, Hua Yang 0001, Shibao Zheng, Yawen Fan, Qin Zhou 0002 |
J. Vis. Commun. Image Represent. | 3 |
| 2017 | Motion sketch based crowd video retrieval
Shuang Wu 0001, Hua Yang 0001, Shibao Zheng, Hang Su 0006, Qin Zhou 0002 |
Multim. Tools Appl. | 2 |
| 2016 | Crowd semantic segmentation based on spatial-temporal dynamicsabstractCrowd semantic segmentation is supposed to not only accurately segment the crowd into groups but also describe them by semantic properties. We define a group as a set of members sharing common spatial-temporal dynamics, i.e., motion consistency and distribution homogeneity. This paper proposes a novel crowd semantic segmentation method, termed as joint spatial-temporal semantic segmentation, which leverages the temporal motion characteristics and spatial distribution information of crowd. We first conduct temporal motion grouping and spatial distribution grouping according to motion consistency and distribution homogeneity respectively. Then, a a joint semantic segmentation algorithm is employed to combine the motion and distribution groups into semantic groups. States of these groups are described in terms of motion pattern and density level. Experiments show that our proposed method is effective to obtain favorable segmentation with semantic descriptions. Jijia Li, Hua Yang 0001, Shuang Wu 0001 |
AVSS | 2 |
| 2016 | Joint instance and feature importance re-weighting for person reidentificationabstractPerson reidentification refers to the task of recognizing the same person under different non-overlapping camera views. Presently, person reidentification based on metric learning is proved to be effective among various techniques, which exploits the labeled data to learn a subspace that maximizes the inter-person divergence while minimizes the intra-person divergence. However, these methods fail to take the different impacts of various instances and local features into account. To address this issue, we propose to learn a projection matrix such that the importance of different instances and local features are re-weighted jointly. We also come up with a simplified formulation of the proposed algorithm, thus it can be solved by the efficient UDFS optimization algorithm. Extensive experiments on the VIPeR and iLIDS datasets demonstrate the effectiveness and efficiency of our algorithm. Qin Zhou 0002, Shibao Zheng, Hua Yang 0001, Hang Su 0006 |
ICASSP | 3 |
| 2016 | Motion sketch based crowd video retrieval via motion structure codingabstractCrowd video retrieval is an important problem in surveillance video management in the era of big data, e.g., video indexing and browsing. In this paper, we address this issue from the motion-level perspective by using hand-drawn sketches as queries. Motion sketch based crowd video retrieval naturally suffers from challenges in motion-level video indexing and sketch representation. We tackle them by leveraging the motion structure coding algorithm to extract robust structure-preserved motion descriptors. For video indexing, we use motion decomposition to separate the sub-motion vector fields with typical patterns from a set of optical flows. Then, the motion-level descriptors of the vector fields are computed and stored in the index database. To represent sketch queries, we propose a sketch vectorization algorithm followed by motion structure coding. In the retrieval stage, given a new query, the retrieval function learned by the Ranking SVM algorithm predicts the ranking score of each motion pattern in the index database. Extensive experiments are conducted on the publicly available crowd datasets, which demonstrate the robustness and effectiveness of the proposed sketch based crowd video retrieval system. Shuang Wu 0001, Hang Su 0006, Shibao Zheng, Hua Yang 0001, Qin Zhou 0002 |
ICIP | 4 |
| 2016 | Person Re-identification Using Cascade Filter
Xinyu Wang 0019, Hua Yang 0001, Ji Zhu 0002, Lin Chen 0019 |
IDEAL | 2 |
| 2016 | Resolution adaptive feature extracting and fusing framework for person re-identification
Hua Yang 0001, Xinyu Wang 0019, Ji Zhu 0002, Wenqi Ma, Hang Su 0006 |
Neurocomputing | 1 |
| 2015 | Kernelized View Adaptive Subspace Learning for Person Re-identification
Qin Zhou 0002, Shibao Zheng, Hang Su 0006, Hua Yang 0001, Shuang Wu 0001 |
BMVC | 4 |
| 2015 | Towards active annotation for detection of numerous and scattered objectsabstractObject detection is an active study area in the field of computer vision and image understanding. In this paper, we propose an active annotation algorithm by addressing the detection of numerous and scattered objects in a view, e.g., hundreds of cells in microscopy images. In particular, object detection is implemented by classifying pixels into specific classes with graph-based semi-supervised learning and grouping neighboring pixels with the same label. Sample or seed selection is conducted based on a novel annotation criterion that minimizes the expected prediction error. The most informative samples are therefore annotated actively, which are subsequently propagated to the unlabeled samples via a pairwise affinity graph. Experimental results conducted on two real world datasets validate that our proposed scheme quickly reaches high quality results and reduces human efforts significantly. Hang Su 0006, Hua Yang 0001, Shibao Zheng, Sha Wei, Shuang Wu 0001 |
ICME | 2 |
| 2015 | Hierarchical video summarization with loitering indicationabstractIn this paper, a hierarchical and informative summarization framework is proposed, which facilitates rapid video browsing. Moreover, a method for loitering detection is exploited to indicate potential abnormal behaviors. The hierarchical framework includes two levels: a holistic-level and an object-level. The holistic-level summarization provides viewers with a comprehensive and compact representation of the original video, while the object-level summarization extracts the narrative information of each object, including trajectory, direction, time, changes of appearance and indication of the loitering behavior. The two summarizations are formulated as two different energy minimization problems, which are solved by the proposed heuristic algorithms. Our framework is evaluated on two publicly available datasets. Experimental results demonstrate that the proposed method performs favourably in providing holistic- and object-level information, fast browsing, and loitering detection. Ruipeng Lu, Hua Yang 0001, Ji Zhu 0002, Shuang Wu 0001, Jia Wang 0004, David Bull 0001 |
VCIP | 2 |
| 2014 | The large-scale crowd analysis based on sparse spatial-temporal local binary pattern
Hua Yang 0001, Yihua Cao, Hang Su 0006, Yawen Fan, Shibao Zheng |
Multim. Tools Appl. | 1 |
| 2013 | Vehicle logo recognition based on Bag-of-WordsabstractThe recognition of vehicle manufacturer logo is a crucial and very challenging problem, which is still an area with few published effective methods. This paper proposes a new fast and reliable system for Vehicle Logo Recognition (VLR) based on Bag-of-Words (BoW). In our system, vehicle logo images are represented as histograms of visual words and classified by SVM in three steps: firstly, extract dense-SIFT features; secondly, quantize features into visual words by `Soft-assignment' thirdly, build histograms of visual words with spatial information. Compared with traditional VLR methods, experiment results show that our proposed system achieves higher recognition accuracy with less processing time. The proposed system is evaluated on a dataset of 840 low-resolution vehicle logo images with about 30×30 pixels, which verifies that our system is practical and effective. Shuyuan Yu, Shibao Zheng, Hua Yang 0001, Longfei Liang |
AVSS | 3 |
| 2013 | The Large-Scale Crowd Behavior Perception Based on Spatio-Temporal Viscous Fluid FieldabstractOver the past decades, a wide attention has been paid to crowd control and management in the intelligent video surveillance area. Among the tasks for automatic surveillance video analysis, crowd motion modeling lays a crucial foundation for numerous subsequent analysis but encounters many unsolved challenges due to occlusions among pedestrians, complicated motion patterns in crowded scenarios, etc. Addressing the unsolved challenges, the authors propose a novel spatio-temporal viscous fluid field to model crowd motion patterns by exploring both appearance of crowd behaviors and interaction among pedestrians. Large-scale crowd events are hereby recognized based on characteristics of the fluid field. First, a spatio-temporal variation matrix is proposed to measure the local fluctuation of video signals in both spatial and temporal domains. After that, eigenvalue analysis is applied on the matrix to extract the principal fluctuations resulting in an abstract fluid field. Interaction force is then explored based on shear force in viscous fluid, incorporating with the fluctuations to characterize motion properties of a crowd. The authors then construct a codebook by clustering neighboring pixels with similar spatio-temporal features, and consequently, crowd behaviors are recognized using the latent Dirichlet allocation model. The convincing results obtained from the experiments on published datasets demonstrate that the proposed method obtains high-quality results for large-scale crowd behavior perception in terms of both robustness and effectiveness. Hang Su 0006, Hua Yang 0001, Shibao Zheng, Yawen Fan, Sha Wei |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2012 | Crowd Event Perception Based on Spatio-temporal Viscous Fluid FieldabstractOver the past decades, a wide attention has been paid to crowd control and management in intelligent video surveillance area. In this paper, the authors propose a novel spatiotemporal viscous fluid field to recognize large-scale crowd event with respect to both appearance and driven factor of crowd behavior. Firstly, a spatiotemporal variation matrix is proposed to exploit motion property of a crowd. In particular, the paper exploits characteristics of the matrix with eigenvalue decomposition algorithm and constructs an abstract fluid field to model the crowd motion pattern, which is denoted by spatiotemporal fluid field. Secondly, the paper proposes a spatiotemporal force field to exploit the interaction force between the pedestrians. Furthermore, the fluid and force field constructs a spatiotemporal viscous fluid field. Thirdly, after generating feature with bag of word model, the authors utilize latent Dirichlet allocation model to recognize crowd behavior. The experiments on PETS2009 and UMN datasets show that the proposed method has a better performance for large-scale crowd behavior perception in both robustness and effectiveness comparing with the conventional methods. Hang Su 0006, Hua Yang 0001, Shibao Zheng, Yawen Fan, Sha Wei |
AVSS | 2 |
| 2011 | Geometric Motion Flow (GMF): A New Feature for Traffic SurveillanceabstractMotion analysis is still a challenging task in many computer vision applications. This paper proposes a new low-level motion feature based on geometric regularity information, particularly suited for traffic surveillance. Firstly, a novel concept of temporal geometry consistency constraint (TGCC) is introduced, which exploits the fact that the geometric structure of a rigid object remains consistent across consecutive frames. Furthermore, the spatial geometric flow is adopted to characterize image structure. Finally, the video motion is represented as a set of geometric flow that moves in the temporal direction. In this case, the method yields a promising illumination robust moderately dense geometric motion flow (GMF) and has more explicit motion boundaries. The GMF could also be used for higher level motion modeling and structural inference tasks, as an effective low-level feature. Extensive experiment results on real video demonstrate the effectiveness and robustness of the proposed method for vehicle motion analysis. Yawen Fan, Hua Yang 0001, Shibao Zheng, Hang Su 0006 |
ICIG | 2 |
| 2011 | The large-scale crowd density estimation based on sparse spatiotemporal local binary patternabstractOver the past decade, a wide attention has been paid to the crowd control and management in intelligent video surveillance area. This paper proposes a sparse spatiotemporal local binary pattern (SST-LBP) descriptor to extract the dynamic texture of the walking crowd with the application to crowd density estimation. Firstly, the sparse selected location is extracted, which is notably variant in temporal domain and scale invariant in spatial domain. Afterwards, considering the spatial and temporal symmetry, the authors propose a sparse spatiotemporal local binary pattern algorithm and utilize its statistical property to describe the crowd feature. Finally, the crowd features are classified into a range of density levels by adopting support vector machine. The experiments on real video show that the proposed SST-LBP method is effective and robust on the large-scale crowd density estimation. Compared with the other methods, the proposed method does not base on the premise that the background should be extracted perfectly, which is too complicated to implement in real surveillance. Hua Yang 0001, Hang Su 0006, Shibao Zheng, Sha Wei, Yawen Fan |
ICME | 1 |
| 2010 | The Large-Scale Crowd Density Estimation Based on Effective Region Feature Extraction Method
Hang Su 0006, Hua Yang 0001, Shibao Zheng |
ACCV (3) | 2 |
| 2009 | Distortion-Minimized Video Slicing for Unequal Loss ProtectionabstractThis letter proposes a distortion-minimized slicing scheme for the unequal loss protection of compressed video bitstreams transported over packet networks. Unlike most existing slicing methods where each slice includes nearly equal number of video macroblocks, the proposed scheme reorders the macroblocks in one video frame according to their importance, and then divides them into two slices with an unequal proportion. According to given channel conditions and a novel unequal loss protection scheme, the more appropriate macroblock division ratio can be found out to achieve minimized end-to-end distortion. Simulation results show that the proposed slicing scheme outperforms state-of-the-art approaches with fixed macroblock division ratio. Hua Yang 0001, Shibao Zheng |
IEEE Signal Process. Lett. | 2 |
| 2008 | GOP-level transmission distortion modeling for mobile streaming video
Hua Yang 0001, Songyu Yu, Xiaokang Yang 0001 |
Signal Process. Image Commun. | 2 |
| 2007 | Cross-Layer Frame Discarding for Cellular Video CodingabstractIn the case of delivering real-time video over the 3G cellular networks, burst frame losses may be inevitable and unpredictable, which may cause severe quality degradation. Based on cross-layer frame discarding (CLFD), this paper proposes an enhanced error-resilient video coding scheme for cellular video communication. By using unequal retransmission at the radio link (RL) layer, a base station can provide reliable transmission for the relatively important frames in one video sequence. Relying on the unequal protection at the RL layer, the encoder at the application (APP) layer can actively discard a certain number of frames according to the received acknowledgement messages. Thus, unpredictable burst frame losses during transmission can be transformed into selective frame discarding at the encoder. Experiments results show that the proposed scheme can enhance the error resilience of the cellular video communication significantly. Songyu Yu, Hua Yang 0001, Hongkai Xiong |
ICASSP (2) | 3 |
| 2007 | GOP-Level Transmission Distortion Modeling for Unequal Importance JudgementabstractIn the case of transmitting stored video streaming over packet-switched networks, unavoidable frame losses may result in error propagation of reconstructed video and thus induce severe quality degradation. In order to minimize the transmission distortion and utilize the limited channel resources efficiently, unequal error protection is usually adopted to exploit the unequal importance of different frames in one group-of-pictures (GOP). In this paper, we develop an estimate model for GOP-level transmission distortion so that the transmitter is able to find the unequal importance of different frames in one GOP. The simulation results demonstrate that the proposed model is accurate and robust. Hua Yang 0001, Songyu Yu, Xiaokang Yang 0001, Hao Liu 0010 |
ICME | 2 |
| 2006 | Motion Vector Smoothing for True Motion EstimationabstractThis paper proposes a new motion vector (MV) smoothing algorithm to track the real motion in image sequences for MPEG video encoders. First, a pre-checking algorithm is employed to eliminate wrong motion vectors and preserve all possible motion vectors. For each block considered, the motion similarity between the neighboring blocks and the number of candidate motion vectors are jointly exploited to adaptively grow the filtering support, which is supposed to have homogeneous motion and sufficient spatial gradient. Then, all candidate motion vectors are checked within the filtering support using a new motion smoothness-constrained matching criteria. The simulation results show that the proposed algorithm can efficiently track the real motion resulting in smooth motion vector field (MVF). Hai Bing Yin, Xiangzhong Fang, Hua Yang 0001, Songyu Yu, Xiaokang Yang 0001 |
ICASSP (2) | 3 |