VLDB 2026 Research / reviewers in the wild / expert
Xin Xu 0007
dblp:66/3874-7
· DBLP profile ↗
90ranked-venue papers
19as first author
55since 2021 · last 2026
0000-0003-0748-3669ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 55 · 11 first-author · 40 since 2021Artificial intelligence and machine learning · 26 · 7 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 4 first-author · 5 since 2021Computer networks · 7 · 2 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-authorDatabases, data management, data science and information retrieval · 3 · 3 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Robust Pedestrian Detection with Uncertain ModalityabstractExisting cross-modal pedestrian detection (CMPD) employs complementary information from RGB and thermal-infrared (TIR) modalities to detect pedestrians in 24h-surveillance systems. RGB captures rich pedestrian details under daylight, while TIR excels at night. However, TIR focuses primarily on the person's silhouette, neglecting critical texture details essential for detection. While the near-infrared (NIR) captures texture under low-light conditions, which effectively alleviates performance issues of RGB and detail loss in TIR, thereby reducing missed detections. To this end, we construct a new Triplet RGB–NIR–TIR (TRNT) dataset, comprising 8,281 pixel-aligned image triplets, establishing a comprehensive foundation for algorithmic research. However, due to the variable nature of real-world scenarios, imaging devices may not always capture all three modalities simultaneously. This results in input data with unpredictable combinations of modal types, which challenge existing CMPD methods that fail to extract robust pedestrian information under arbitrary input combinations, leading to significant performance degradation. To address these challenges, we propose the Adaptive Uncertainty-aware Network (AUNet) for accurately discriminating modal availability and fully utilizing the available information under uncertain inputs. Specifically, we introduce Unified Modality Validation Refinement (UMVR), which includes an uncertainty-aware router to validate modal availability and a semantic refinement to ensure the reliability of information within the modality. Furthermore, we design a Modality-Aware Interaction (MAI) module to adaptively activate or deactivate its internal interaction mechanisms per UMVR output, enabling effective complementary information fusion from available modalities. AUNet enables accurate modality validation and robust inference without fixed modality pairings, facilitating the effective fusion of RGB, NIR, and TIR information across diverse inputs. Qian Bie, Xiao Wang 0029, Bin Yang 0026, Zhixi Yu, Jun Chen 0001, Xin Xu 0007 |
AAAI | 6 |
| 2026 | MambaOVSR: Multiscale Fusion with Global Motion Modeling for Chinese Opera Video Super-ResolutionabstractChinese opera is celebrated for preserving classical art. However, early filming equipment limitations have degraded videos of last-century performances by renowned artists (e.g., low frame rates and resolution), hindering archival efforts. Although space-time video super-resolution (STVSR) has advanced significantly, applying it directly to opera videos remains challenging. The scarcity of datasets impedes the recovery of high-frequency details, and existing STVSR methods lack global modeling capabilities—compromising visual quality when handling opera’s characteristic large motions. To address these challenges, we pioneer a large-scale Chinese Opera Video Clip (COVC) dataset and propose the Mamba-based multiscale fusion network for space-time Opera Video Super-Resolution (MambaOVSR). Specifically, MambaOVSR involves three novel components: the Global Fusion Module (GFM) for motion modeling through a multiscale alternating scanning mechanism, and the Multiscale Synergistic Mamba Module (MSMM) for alignment across different sequence lengths. Additionally, our MambaVR block resolves feature artifacts and positional information loss during alignment. Experimental results on the COVC dataset show that MambaOVSR significantly outperforms the SOTA STVSR method by an average of 1.86 dB in terms of PSNR. Hua Chang, Xin Xu 0007, Wei Liu 0183, Wei Wang 0170, Xin Yuan 0009, Kui Jiang |
AAAI | 2 |
| 2026 | ICLR: Inter-Chrominance and Luminance Interaction for Natural Color Restoration in Low-Light Image EnhancementabstractLow-Light Image Enhancement (LLIE) task aims at improving contrast while restoring details and textures for images captured in low-light conditions. HVI color space has made significant progress in this task by enabling precise decoupling of chrominance and luminance. However, for the interaction of chrominance and luminance branches, substantial distributional differences between the two branches prevalent in natural images limit complementary feature extraction, and luminance errors are propagated to chrominance channels through the nonlinear parameter. Furthermore, for interaction between different chrominance branches, images with large homogeneous-color regions usually exhibit weak correlation between chrominance branches due to concentrated distributions. Traditional pixel-wise losses exploit strong inter-branch correlations for co-optimization, causing gradient conflicts in weakly correlated regions. Therefore, we propose an Inter-Chrominance and Luminance Interaction (ICLR) framework including a Dual-stream Interaction Enhancement Module (DIEM) and a Covariance Correction Loss (CCL). The DIEM improves the extraction of complementary information from two dimensions, fusion and enhancement, respectively. The CCL utilizes luminance residual statistics to penalize chrominance errors and balances gradient conflicts by constraining chrominance branches covariance. Experimental results on multiple datasets show that the proposed ICLR framework outperforms state-of-the-art methods. Xin Xu 0007, Wei Liu 0183, Wei Wang 0170, Kui Jiang |
AAAI | 1 |
| 2026 | FedARKS: Federated Aggregation via Robust and Discriminative Knowledge Selection and Integration for Person Re-identification
Xin Xu 0007, Binchang Ma, Zhixi Yu, Wei Liu 0183 |
AAAI | 1 |
| 2026 | Domain-Aware Suppression and Aggregation for Federated DG ReIDabstractFederated domain generalization in person re-identification (FedDG-ReID) aims to learn a privacy-preserving server model from decentralized client source domains that generalizes to unseen domains. Existing approaches enhance the generalizability of the server model by increasing the diversity of client person data. However, these methods overlook that ReID model parameters are easily biased by client-specific data distributions, leading to the capture of excessive domain-specific identity information. Such identity information (e.g., clothing style) struggles with identity information in unseen domains, thereby hindering the generalization ability of the server model. To address this, we propose a novel FedDG-ReID framework, which mainly consists of Domain-aware Parameter Suppression (DPS) and Domain-invariant Weighted Aggregation (DWA), called FedSupWA. Specifically, DPS adaptively attenuates the update magnitude of the parameters based on the fit of the parameters to the client's domain, encouraging the model to focus on more generalized domain-independent identity information, such as pedestrian contours, and other consistent information across domains. DWA enhances the server model’s generalization by evaluating the effectiveness of the client model in maintaining the consistency of pedestrian identities to measure the importance of the learned domain-independent identity information and assigning greater aggregation weights to clients that contribute more generalized information. Extensive experiments demonstrate the effectiveness of FedSupWA, showing that it achieves state-of-the-art performance. Zhixi Yu, Wei Liu 0183, Wenke Huang 0003, Bin Yang 0026, Qian Bie, Guancheng Wan, Xin Xu 0007 |
AAAI | 7 |
| 2026 | NPFML: Non-isotropic Potential Fields with Hierarchical Decay for Deep Metric Learning
Xin Yuan 0009, Minshi Chen, Xin Xu 0007 |
MMM (2) | 5 |
| 2026 | DSFMamba: Dual-state fusion mamba for long-short term motion-adaptive streaming perception in autonomous driving
Qinghua Qi, Xiaowei Xu 0003, Mingxing Deng, Wei Liu 0183, Xin Xu 0007 |
Knowl. Based Syst. | 5 |
| 2026 | Discrepancy-Consistency Bi-Knowledge Fusion for unsupervised video anomaly detection
Awei Yin, Wei Liu 0183, Xiao Wang 0029, Ao Huang, Xin Xu 0007 |
Knowl. Based Syst. | 6 |
| 2026 | TH-Mamba: Spatial-Temporal Correlation Learning for Mamba-Based Talking Head GenerationabstractTalking head generation aims to synthesize high-quality and lip-synchronized talking head videos from the given portrait images and audio. However, previous methods directly learn the alignment between lip movements and the driven audio, barely focusing on the fidelity and continuity of the generated videos, suffering from visual distortions and jitter. To deal with this issue, we propose to promote the consistency of audio and image by exploring their spatiotemporal relations, and construct a Mamba-based spatiotemporal fusion scheme. Specifically, we devise an Intra-frame Mamba module to characterize facial features from the source image, which encourages the content consistence between the generated frame and the current source frame. Meanwhile, an Inter-frame Mamba module is designed to excavate the complementary information across sequential frames, which provides clues for better motion simulation. The aggregated spatiotemporal representation with audio features are then aligned with a deformation network to alleviate visual distortions and jitter. In addition, we investigate the practical composite constraints on the structure, details, and motion aspects, involving the keypoint constraint, multi-scale content constraint, and displacement constraint to promote the training stability and model performance. With the above strategies, we construct a novel Talking Head Mamba network, termed as TH-Mamba for high-quality talking head generation. Extensive experiments on the HDTF and Mead-Neutral datasets verify the superiority of our proposed TH-Mamba, which significantly outperforms the current state-of-the-art method by 0.78dB and 1.55dB in PSNR, respectively. The demo is available at https://github.com/YZX-codesky/TH-Mamba. Xin Xu 0007, Zhixi Yu, Kui Jiang, Chia-Wen Lin |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | TokenMatcher: Diverse Tokens Matching for Unsupervised Visible-Infrared Person Re-IdentificationabstractUnsupervised visible-infrared person re-identification (US-VI-ReID) seeks to match infrared and visible images of the same individual without the use of annotations. Current methods typically derive cross-modal correspondences through a single global feature matching process for generating pseudo labels and learning modality-invariant features. However, this matching approach is hindered by both intra-modality and inter-modality discrepancies, which result in imprecise measurements. As a consequence, the clustering of individuals with single global feature is often incomplete and unreliable, leading to suboptimal performance in cross-modal clustering tasks. To address these challenges and to extract cross-modality discriminative identity information, we propose a TokenMatcher, which encompasses three key components: Diverse Tokens Matching (DTM), Diverse Tokens Neighbor Learning (DTNL), and the Homogeneous Fusion (HF) Module. DTM utilizes multiple class tokens within the visual transformer framework to capture diverse embedding representations, thereby facilitating the integration of fine-grained information essential for reliable cross-modality correspondences. DTNL enhances the intra-modality and inter-modality consistency among diverse tokens by refining neighborhood sets with insights from neighboring tokens and camera information, promoting robust neighborhood learning and fostering discriminative identity information. Additionally, the HF module consolidates clusters of the same identity while effectively separating those of different identities. Extensive experiments conducted on the publicly available SYSU-MM01 and RegDB datasets demonstrate the efficacy of the proposed method. Xiao Wang 0029, Lekai Liu, Bin Yang 0026, Mang Ye, Zheng Wang 0007, Xin Xu 0007 |
AAAI | 6 |
| 2025 | The Parables of the Mustard Seed and the Yeast: Extremely Low-Budget, High-Performance Nighttime Semantic SegmentationabstractNighttime Semantic Segmentation (NSS) is essential to many cutting-edge vision applications. However, existing technologies overly rely on massive labeled data, whose annotation is time-consuming and laborious. In this paper, we pioneer a new task focusing on exploring the potential of training strategy and framework design with limited annotation to achieve high-performance NSS. Insufficient information at very low labeling budgets can easily lead to under-optimization or overfitting of the model. Our solution comprises two main components: i) a novel region-based active sampling strategy called Contextual-Aware Region Query (CARQ), which identifies highly informative target nighttime regions for labeling; and ii) an innovative Fragmentation Synergy Active Domain Adaptation framework (FS-ADA), which progressively broadcasts the limited annotation to the unlabeled regions, achieving high performance with a minimal annotation budget. Extensive experiments demonstrate that our method outperforms state-of-the-art UDA-NSS & ADA-SS methods across four day-to-nighttime benchmarks, and generalizes well to foggy, rainy, & snowy scenes. In particular only with 1% target nighttime data annotation, our method is on par with the mainstream fully-supervised methods on the BDD100K-Night val dataset. Shiqin Wang, Xin Xu 0007, Kui Jiang, Zheng Wang 0007 |
AAAI | 2 |
| 2025 | VAGeo: View-specific Attention for Cross-View Object Geo-LocalizationabstractCross-view object geo-localization (CVOGL) aims to locate an object of interest in a captured ground- or drone-view image within the satellite image. However, existing works treat ground-view and drone-view query images equivalently, overlooking their inherent viewpoint discrepancies and the spatial correlation between the query image and the satellite-view reference image. To this end, this paper proposes a novel View-specific Attention Geo-localization method (VAGeo) for accurate CVOGL. Specifically, VAGeo contains two key modules: view-specific positional encoding (VSPE) module and channel-spatial hybrid attention (CSHA) module. In object-level, according to the characteristics of different viewpoints of ground and drone query images, viewpoint-specific positional codings are designed to more accurately identify the click-point object of the query image in the VSPE module. In feature-level, a hybrid attention in the CSHA module is introduced by combining channel attention and spatial attention mechanisms simultaneously for learning discriminative features. Extensive experimental results demonstrate that the proposed VAGeo gains a significant performance improvement, i.e., improving [email protected]/[email protected] on the CVOGL dataset from 45.43%/42.24% to 48.21%/45.22% for ground-view, and from 61.97%/57.66% to 66.19%/61.87% for drone-view. Xin Yuan 0009, Wei Liu 0183, Xin Xu 0007 |
ICASSP | 4 |
| 2025 | Event-based Video Person Re-identification via Cross-Modality and Temporal CollaborationabstractVideo-based person re-identification (ReID) has become increasingly important due to its applications in video surveillance applications. By employing events in video-based person ReID, more motion information can be provided between continuous frames to improve recognition accuracy. Previous approaches have assisted by introducing event data into the video person ReID task, but they still cannot avoid the privacy leakage problem caused by RGB images. In order to avoid privacy attacks and to take advantage of the benefits of event data, we consider using only event data. To make full use of the information in the event stream, we propose a Cross-Modality and Temporal Collaboration (CMTC) network for event-based video person ReID. First, we design an event transform network to obtain corresponding auxiliary information from the input of raw events. Additionally, we propose a differential modality collaboration module to balance the roles of events and auxiliaries to achieve complementary effects. Furthermore, we introduce a temporal collaboration module to exploit motion information and appearance cues. Experimental results demonstrate that our method outperforms others in the task of event-based video person ReID. Renkai Li, Xin Yuan 0009, Wei Liu 0183, Xin Xu 0007 |
ICASSP | 4 |
| 2025 | Foundation Model Insights and a Multi-Model Approach for Superior Fine-Grained One-shot Subset SelectionabstractOne-shot subset selection serves as an effective tool to reduce deep learning training costs by identifying an informative data subset based on the information extracted by an information extractor (IE). Traditional IEs, typically pre-trained on the target dataset, are inherently dataset-dependent. Foundation models (FMs) offer a promising alternative, potentially mitigating this limitation. This work investigates two key questions: (1) Can FM-based subset selection outperform traditional IE-based methods across diverse datasets? (2) Do all FMs perform equally well as IEs for subset selection? Extensive experiments uncovered surprising insights: FMs consistently outperform traditional IEs on fine-grained datasets, whereas their advantage diminishes on coarse-grained datasets with noisy labels. Motivated by these finding, we propose RAM-APL (RAnking Mean-Accuracy of Pseudo-class Labels), a method tailored for fine-grained image datasets. RAM-APL leverages multiple FMs to enhance subset selection by exploiting their complementary strengths. Our approach achieves state-of-the-art performance on fine-grained datasets, including Oxford-IIIT Pet, Food-101, and Caltech-UCSD Birds-200-2011. Zhijing Wan, Zhixiang Wang 0001, Zheng Wang 0007, Xin Xu 0007, Shin'ichi Satoh 0001 |
ICML | 4 |
| 2025 | Positive Style Accumulation: A Style Screening and Continuous Utilization Framework for Federated DG-ReIDabstractThe Federated Domain Generalization for Person re-identification (FedDG-ReID) aims to learn a global server model that can be effectively generalized to source and target domains through distributed source domain data. Existing methods mainly improve the diversity of samples through style transformation, which to some extent enhances the generalization performance of the model. However, we discover that not all styles contribute to the generalization performance. Therefore, we define styles that are beneficial/harmful to the model's generalization performance as positive/negative styles. Based on this, new issues arise: How to effectively screen and continuously utilize the positive styles. To solve these problems, we propose a Style Screening and Continuous Utilization (SSCU) framework. Firstly, we design a Generalization Gain-guided Dynamic Style Memory (GGDSM) for each client model to screen and accumulate generated positive styles. Specifically, the memory maintains a prototype initialized from raw data for each category, then screens positive styles that enhance the global model during training, and updates these positive styles into the memory using a momentum-based approach. Meanwhile, we propose a style memory recognition loss to fully leverage the positive styles memorized by GGDSM. Furthermore, we propose a Collaborative Style Training (CST) strategy to make full use of positive styles. Unlike traditional learning strategies, our approach leverages both newly generated styles and the accumulated positive styles stored in memory to train client models on two distinct branches. This training strategy is designed to effectively promote the rapid acquisition of new styles by the client models, ensuring that they can quickly adapt to and integrate novel stylistic variations. Simultaneously, this strategy guarantees the continuous and thorough utilization of positive styles, which is highly beneficial for the model's generalization performance. Extensive experimental results demonstrate that our method outperforms existing methods in both the source domain and the target domain. Xin Xu 0007, Chaoyue Ren, Wei Liu 0183, Wenke Huang 0003, Bin Yang 0026, Zhixi Yu, Kui Jiang |
ACM Multimedia | 1 |
| 2025 | Style Separation and Content Recovery for Generalizable Sketch Re-identification and a New Benchmark
Lingyi Lu, Xin Xu 0007, Xiao Wang 0029 |
MMM (4) | 2 |
| 2025 | Retrieving and Reasoning: Multivariate Feature and Attribute Cooperation for Video Anomaly DetectionabstractVideo anomaly detection (VAD), which detects abnormal patterns in video sequence, is based on several kinds of features or attributes in the existing methods. This ignores the interconnections between different features and attributes, and the initiation of an anomalous result is brought about by multiple factors. If several individual neural networks are used to perceive various types of anomalies, the system would lose awareness of the association among features and attributes, which limits the system's ability to perceive complex anomalies. In this work, we propose a dual-branch framework for VAD task, which includes deep feature retrieving and semantic attribute reasoning branch. In the former branch, three high-dimensional deep features are extracted and modeled, then the anomaly scores are obtained based on the vector retrieval database. In the latter branch, three low-dimensional semantic-level attributes are extracted for composing the attribute triplets, then use theAssociation-ruleMiningModule (AMM) to perceive potential connections among these triplets. The coefficients computed by the latter branch calibrate the anomaly scores obtained by the former while providing high-level anomaly causes. Extensive experiments show that our approach achieves state-of-the-art performance with 87.9$\%$on ShanghaiTech and 94.6$\%$on Avenue. Xingshuo Han, Xiao Wang 0029, Wei Liu 0183, Liping Ye, Xin Xu 0007 |
IEEE Signal Process. Lett. | 5 |
| 2025 | Multi-Objective Optimization for Multimodal Multi-Objective Multi-Point Shortest Path Problem Considering Unforeseeable Road EventualitiesabstractMulti-objective multi-point shortest path planning problems are commonly encountered in real-world applications. Numerous path planning algorithms have been proposed to accommodate different model assumptions. However, most existing algorithms can only identify a subset of the Pareto optimal paths and overlook equivalent Pareto optimal paths. Relying solely on a subset of Pareto optimal solutions is insufficient to effectively respond to unforeseeable road eventualities in the real-world traffic environment. In this paper, multi-objective multi-point shortest path planning problem is modeled as a multimodal multi-objective optimization problem with necessary points constrains. A multimodal multi-objective evolutionary algorithm using constraint dominance principle-based path comparison strategy and path similarity-based multimodal solutions selection strategy is proposed to address this problem. The proposed constraint dominance principle-based path comparison strategy can effectively navigate through large infeasible regions by relaxing necessary point constraints, thereby obtaining a true constrained Pareto front. The proposed path similarity-based multimodal solutions selection strategy can effectively balance the distribution of solutions in the decision space, thereby preserving multiple equivalent optimal solutions. The proposed algorithm is compared with five state-of-the-art path planning algorithms from the benchmark test suite derived from the 2021 IEEE CEC path planning competition, where city maps are adapted from real transportation networks in Chinese cities, in our experiments. The exceptional performance is demonstrated through thirty independent runs, yielding experimental results that showcase the superiority of the proposed algorithm on the test problem set. This superior performance highlights the potential for designing more resilient path planners suitable for scenarios affected by unpredictable road eventualities. Zhiwei Xu 0004, Kai Zhang 0002, Javier Del Ser, Miqing Li, Xin Xu 0007, Juanjuan He, Ni Wu |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | Diversity-Representativeness Replay and Knowledge Alignment for Lifelong Vehicle Re-identificationabstractLifelong Vehicle Re-Identification (LVReID) aims to match a target vehicle across multiple cameras, considering non-stationary and continuous data streams, which fits the needs of the practical application better than traditional vehicle re-identification. Nonetheless, this area has received relatively little attention. Recently, methods for Lifelong Person Re-Identification (LPReID) have been emerging, with replay-based methods achieving the best results by storing a small number of instances from previous tasks for retraining, thus effectively reducing catastrophic forgetting. However, these methods cannot be directly applied to LVReID because they fail to simultaneously consider the diversity and representativeness of replayed data, resulting in biases between the subset stored in the memory buffer and the original data. They randomly sample classes, which may not adequately represent the distribution of the original data. Additionally, these methods fail to consider the rich variation in instances of the same vehicle class due to factors such as vehicle orientation and lighting conditions. Therefore, preserving more informative classes and instances for replay helps maintain information from previous tasks and may mitigate the model's forgetting of old knowledge. In view of this, we propose a novel Diversity-Representativeness Dual-Stage Sampling Replay (DDSR) strategy for LVReID that constructs an effective memory buffer through two stages, i.e. , Cluster-Centric Class Selection and Diverse Instance Mining. Specifically, we first perform class-level sampling based on density in the clustered class-centered feature space and then further mine the diverse, high-quality instances within the selected classes. In addition, we introduce Maximum Mean Discrepancy loss to align the feature distribution between replay data and the new arrivals and apply L2 regularization in the parameter space to facilitate knowledge transfer, thus enhancing the model's generalization ability to new tasks. Extensive experiments demonstrate effective improvements of our method compared to current state-of-the-art lifelong ReID methods on the VeRi-776, VehicleID, and VERI-Wild datasets. Zhijing Wan, Xiao Wang 0029, Wei Liu 0183, Wei Wang 0170, Zheng Wang 0059, Xin Xu 0007 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2025 | Mix-Modality Person Re-Identification: A New and Practical ParadigmabstractCurrent visible-infrared cross-modality person re-identification research has only focused on exploring the bi-modality mutual retrieval paradigm, and we propose a new and more practical mix-modality retrieval paradigm. Existing Visible-Infrared Person Re-Identification (VI-ReID) methods have achieved some results in the bi-modality mutual retrieval paradigm by learning the correspondence between visible and infrared modalities. However, significant performance degradation occurs due to the modality confusion problem when these methods are applied to the new mix-modality paradigm. Therefore, this article proposes a Mix-Modality Person Re-Identification (MM-ReID) task, explores the influence of modality mixing ratio on performance, and constructs mix-modality test sets for existing datasets according to the new mix-modality testing paradigm. To solve the modality confusion problem in MM-ReID, we propose a Cross-Identity Discrimination Harmonization Loss (CIDHL) adjusting the distribution of samples in the hyperspherical feature space, pulling the centers of samples with the same identity closer, and pushing away the centers of samples with different identities while aggregating samples with the same modality and the same identity. Furthermore, we propose a Modality Bridge Similarity Optimization Strategy (MBSOS) to optimize the cross-modality similarity between the query and queried samples with the help of the similar bridge sample in the gallery. Extensive experiments demonstrate that compared to the original performance of existing cross-modality methods on MM-ReID, the addition of our CIDHL and MBSOS demonstrates a general improvement. Wei Liu 0183, Xin Xu 0007, Hua Chang, Xin Yuan 0009, Zheng Wang 0007 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | Glare countering and exploiting via dual stream network for nighttime vehicle detection
Pengshu Du, Xiao Wang 0029, WeiGang Li, Xin Xu 0007 |
Vis. Comput. | 6 |
| 2024 | FMRNet: Image Deraining via Frequency Mutual RevisionabstractThe wavelet transform has emerged as a powerful tool in deciphering structural information within images. And now, the latest research suggests that combining the prowess of wavelet transform with neural networks can lead to unparalleled image deraining results. By harnessing the strengths of both the spatial domain and frequency space, this innovative approach is poised to revolutionize the field of image processing. The fascinating challenge of developing a comprehensive framework that takes into account the intrinsic frequency property and the correlation between rain residue and background is yet to be fully explored. In this work, we propose to investigate the potential relationships among rain-free and residue components at the frequency domain, forming a frequency mutual revision network (FMRNet) for image deraining. Specifically, we explore the mutual representation of rain residue and background components at frequency domain, so as to better separate the rain layer from clean background while preserving structural textures of the degraded images. Meanwhile, the rain distribution prediction from the low-frequency coefficient, which can be seen as the degradation prior is used to refine the separation of rain residue and background components. Inversely, the updated rain residue is used to benefit the low-frequency rain distribution prediction, forming the multi-layer mutual learning. Extensive experiments demonstrate that our proposed FMRNet delivers significant performance gains for seven datasets on image deraining task, surpassing the state-of-the-art method ELFormer by 1.14 dB in PSNR on the Rain100L dataset, while with similar computation cost. Code and retrained models are available at https://github.com/kuijiang94/FMRNet. Kui Jiang, Junjun Jiang, Xianming Liu 0005, Xin Xu 0007, Xianzheng Ma |
AAAI | 4 |
| 2024 | Mutuality Attribute Makes Better Video Anomaly DetectionabstractVideo anomaly detection (VAD) is an essential but challenging task. Existing prevalent methods focus on analyzing the reconstruction or prediction difference between normal and abnormal patterns through multiple deep features, e.g., optic flow. However, these approaches independently use deep features to characterize attributes, ignore the mutuality among multiple deep features. Therefore, the constructed representation is limited to indirectly representing the anomaly from isolated attributes, and makes the network difficult to capture the high-level causes of anomaly. In this paper, we proposed a novel Mutuality Attribute-based Representation framework (MAR-VAD) for the VAD task, which absorbs the mutuality among deep features to characterize the mutuality attribute. Specifically, the mutuality attribute encapsulates high-level semantic information, such as the specific abnormal object or action, which mutually utilizes information from multiple deep features. In this way, the system is able to directly capture the high-level causes of anomaly, thus providing a more comprehensive perspective to accurately detect anomaly events. Following a process-transparent density estimation, we produce the final anomaly scores. Experiments show that MAR-VAD achieves state-of-the-art performance on ShanghaiTech and Avenue. Xingshuo Han, Xiao Wang 0029, Kui Jiang, Wei Liu 0183, Ruimin Hu, Xuefeng Pan, Xin Xu 0007 |
ICASSP | 7 |
| 2024 | A comprehensive survey of visible infrared person re-identification from an application perspective
Hua Chang, Xin Xu 0007, Wei Liu 0183, Lingyi Lu, Weigang Li 0004 |
Multim. Tools Appl. | 2 |
| 2024 | Contour Counts: Restricting Deformation for Accurate Animation InterpolationabstractAnimated videos with low frame rates commonly degrade the visual experience with choppy motion. Recently, video frame interpolation has been growing rapidly which can increase frame rates. However, the sparse texture information and complex motion scenes of animated videos make the objects in the frames generated by existing video frame interpolation methods appear significantly deformation, causing distortion in the content of generated frames. To address this issue, the Restricting Deformation by Contour Network (RDC-Net) is proposed to repair and fill the content by optimizing object contours and leveraging context features for high-quality animation interpolation. Specifically, the RDC-Net was proposed to learn optical flow maps utilized to effectively capture the spatial shifts in the motion subject's contours across time intervals. Furthermore, contour information is used to refine the structure of the object in optical flow estimation and moderate the object deformation scale in the generated frame. In addition, the context information is explored to characterize the motion between adjacent frames via bidirectional optical flow learning, enabling the filling of distortion in the content of generated frames by feature-filling technology. Experiments on commonly used benchmarks show our state-of-the-art performance. Xin Xu 0007, Kui Jiang, Wei Liu 0183, Zheng Wang 0007 |
IEEE Signal Process. Lett. | 2 |
| 2024 | From Multi-Source Virtual to Real: Effective Virtual Data Search for Vehicle Re-IdentificationabstractWithout tedious and time-consuming labeling processes, virtual datasets have recently shown their superiority for vehicle re-identification (re-ID). Existing virtual to real vehicle re-ID methods employ only a single virtual dataset for model training, while datasets from different generative sources are not jointly exploited. Multiple source virtual datasets contain more data diversity that can boost model performance. We thus propose a multi-source virtual to real vehicle re-ID pipeline, where multiple source virtual datasets are used during training. However, the multi-source virtual dataset suffers from more data redundancy than the single virtual dataset, which can affect the training efficiency. Intuitively, it can be mitigated by virtual data search. Unlike a single virtual dataset, a performance gap exists between multiple source virtual datasets, indicating their different contributions to model learning. Accordingly, we propose to split the multi-source virtual dataset into the main training set and the auxiliary training set, and then design the sampling strategy separately. For the main training set, the Consistent Attribute Distribution-FEature distance Trade-off (CAD-FET) strategy is designed to search for representative data. For the auxiliary training set, a cluster-based sampling strategy is further proposed to search for the most diverse subset. Besides, a simple yet effective two-stage training strategy is proposed to utilize these subsets reasonably. Extensive virtual-to-real vehicle re-ID experiments show that our data sampling method can reduce the volume of the multi-source virtual dataset by around 77%/96% and boost the model performance when tested on the VeRi776/VehicleID. Zhijing Wan, Xin Xu 0007, Zheng Wang 0007, Zhixiang Wang 0001, Ruimin Hu |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | Blind 3D Video Stabilization with Spatio-Temporally Varying Motion BlurabstractVideo stabilization is a challenging task that attempts to compensate for the overall frame shake during video acquisition. Existing three-dimensional video stabilization methods aim at modeling camera perspective projection through either data-driven training or explicit motion estimation. However, the above methods are difficult to effectively solve the issue of shaky videos with abrupt object movements, resulting in local motion blur in the direction of the movement. This phenomenon is prevalent in real-world scenarios featuring foreground blind motion scenes. Unfortunately, directly combining stabilization and deblurring methods poses challenges when dealing with this situation. In the video, the intensity of motion blur undergoes continuous changes, and the direct combination method inadequately utilizes spatiotemporal information, providing insufficient clues for cross-frame compensation. To alleviate this problem, the Cross-frame-temporal Module framework is proposed to address blind motion blur induced by various conditions, which utilizes cross-frame temporal features to estimate depth maps and camera motion. In this framework, a Blur Transform Network (BTNet) is designed to adapt to spatially varying motion blur, which transforms local regions according to the impact of blur intensities to adapt to the effects of non-uniform motion blur; furthermore, our Temporal-Aware Network (TANet) further suppresses motion blur by leveraging cross-frame temporal features. In addition, the limited availability of pair-training video data containing motion blur limits the application of this approach in practice. The Cross-frame-temporal Module framework adopts an un-pretrained in-test training strategy. Extensive experimental results have demonstrated that our method outperforms state-of-the-art methods. Hengwei Li, Wei Wang 0170, Xiao Wang 0029, Xin Yuan 0009, Xin Xu 0007 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | Lightweight Separable Convolutional Dehazing Network to Mobile FPGA
Xinrui Ju, Wei Wang 0170, Xin Xu 0007 |
CGI (4) | 3 |
| 2023 | From Generation to Suppression: Towards Effective Irregular Glow Removal for Nighttime Visibility EnhancementabstractMost existing Low-Light Image Enhancement (LLIE) methods are primarily designed to improve brightness in dark regions, which suffer from severe degradation in nighttime images. However, these methods have limited exploration in another major visibility damage, the glow effects in real night scenes. Glow effects are inevitable in the presence of artificial light sources and cause further diffused blurring when directly enhanced. To settle this issue, we innovatively consider the glow suppression task as learning physical glow generation via multiple scattering estimation according to the Atmospheric Point Spread Function (APSF). In response to the challenges posed by uneven glow intensity and varying source shapes, an APSF-based Nighttime Imaging Model with Near-field Light Sources (NIM-NLS) is specifically derived to design a scalable Light-aware Blind Deconvolution Network (LBDN). The glow-suppressed result is then brightened via a Retinex-based Enhancement Module (REM). Remarkably, the proposed glow suppression method is based on zero-shot learning and does not rely on any paired or unpaired training data. Empirical evaluations demonstrate the effectiveness of the proposed method in both glow suppression and low-light enhancement tasks. Wanyu Wu, Wei Wang 0170, Zheng Wang 0007, Kui Jiang, Xin Xu 0007 |
IJCAI | 5 |
| 2023 | Informative Classes Matter: Towards Unsupervised Domain Adaptive Nighttime Semantic SegmentationabstractUnsupervised Domain Adaptive Nighttime Semantic Segmentation (UDA-NSS) aims to adapt a robust model from a labeled daytime domain to an unlabeled nighttime domain. However, current advanced segmentation methods ignore the illumination effect and class discrepancies of different semantic classes during domain adaptation, showing an uneven prediction phenomenon. It is the completely ignored and underexplored issues of ''hard-to-adapt'' classes that some classes have a large performance gap between existing UDA-NSS methods and supervised learning counterparts while others have a very low performance gap. To realize ''hard-to-adapt'' classes' more sufficient learning and facilitate the UDA-NSS task, we present an Online Informative Class Sampling (OICS) strategy to adaptively mine informative classes from the target nighttime domain according to the corresponding spectrogram mean and the class frequency via our Informative Mixture of Experts. Furthermore, an Informativeness-based cross-domain Mixed Sampling (InforMS) framework is designed to focus on informative classes from the target nighttime domain by vesting their higher sampling probabilities when cross-domain mixing sampling and achieves better performance in UDA-NSS tasks. Consequently, our method outperforms state-of-the-art UDA-NSS methods by large margins on three widely-used benchmarks (e.g., ACDC, Dark Zurich, and Nighttime Driving). Notably, our method achieves state-of-the-art performance with 65.1% mIoU on ACDC-night-test and 55.4% mIoU on ACDC-night-val. Shiqin Wang, Xin Xu 0007, Xianzheng Ma, Kui Jiang, Zheng Wang 0007 |
ACM Multimedia | 2 |
| 2023 | Dual-focus: person search from Coarse-Grained Focus to Fine-Grained Focus
Wenyi Hu, Xiao Wang 0029, Zheng Wang 0007, Xin Xu 0007, Ruimin Hu |
Multim. Syst. | 4 |
| 2023 | Active learning with sampling by joint global-local uncertainty for salient object detection
Haidong Fu, Xin Xu 0007 |
Neural Comput. Appl. | 3 |
| 2023 | Lightweight CNN-Based Low-Light-Image Enhancement System on FPGA Platform
Wei Wang 0170, Xin Xu 0007 |
Neural Process. Lett. | 2 |
| 2023 | From Collective Attribute Association of Groups to Precise Attribute Association of IndividualsabstractObscured person re-identification (Re-ID) aims to match an obscured image with a complete image of the same person captured by other cameras. As a major challenge in person identification, occlusion severely affects the effectiveness of most traditional person Re-ID methods. To solve this problem, this study proposes a trajectory association method, which, as a pre-processing technique for person Re-ID, can narrow the search range and reduce the problem of degradation caused by mixing. We investigate the method of converting the fuzzy association between sets into the precise association between elements for M video objects and N phone objects (trajectory information) with fuzzy group association relationships at the crime scene. First, we decompose the M-N precise association problem and analyze the similarity of the video objects in the source point and on the trajectories. Then, we define high-similarity points, study their distribution characteristics in different trajectories, and find that there is a significant difference between the distribution of high-similarity points in correct and incorrect matching trajectories. We simplify the full-path association problem into a partial-path high-similarity point distribution difference problem, which effectively reduces the difficulty in accurate association relationship construction. The association experiments in simple and mixed scenarios as well as Re-ID experiments on the PRPW and Market1501 demonstrate the effectiveness of our method. Yun Lan, Ruimin Hu, Xin Xu 0007, Dengshi Li, Chao Wang 0084, Xiaochen Wang 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | Low-light image enhancement with joint illumination and noise data distribution transformation
Wei Wang 0170, Xiao Wang 0029, Xin Xu 0007 |
Vis. Comput. | 4 |
| 2022 | Self-Supervised Learning on A Lightweight Low-Light Image Enhancement Model with Curve RefinementabstractDeep learning networks with deeper layers become a trend for their good performance but lacks the potential for real-time mobile deployment. Another challenge for paired training networks is the limited generalization capacity caused by the sample bias. To overcome these two challenges, we propose a lightweight self-supervised low-light image enhancement method, that trains with low light images only. Specifically, our method consists of a low-resolution dense CNN network stream and a full-resolution guidance stream, responsible for image-to-curve transformation with refinement and spatial guidance fusion, respectively. Then, a new self-supervised loss function is introduced to measure the restored patch-based color deviations among color channels. Experimental results show that our method gives competitive performance to the full-supervised approaches. Wanyu Wu, Wei Wang 0170, Kui Jiang, Xin Xu 0007, Ruimin Hu |
ICASSP | 4 |
| 2022 | Towards Causality Inference for Very Important Person LocalizationabstractVery Important Person Localization (VIPLoc) aims at detecting certain individuals in a given image, who are more attractive than others in the image. Existing uncontrolled VIPLoc benchmark assumes that the image has one single VIP, which is not suitable for actual application scenarios when multiple VIPs or no VIPs appear in the image. In this paper, we re-built a complex uncontrolled conditions (CUC) dataset to make the VIPLoc closer to the actual situation, containing no, single, and multiple VIPs. Existing methods use the hand-designed and deep learning strategies to extract the features of persons and analyze the differences between VIPs and other persons from the perspective of statistics. They are not explainable as to why the VIP located this output for that input. Thus, there exist the severe performance degradation when we use these models in real-world VIPLoc. Specifically, we establish a causal inference framework that unpacks the causes of previous methods and derives a new principled solution for VIPLoc. It treats the scene as confounding factor, allowing the ever-elusive confounding effects to be eliminated and the essential determinants to be uncovered. Through extensive experiments, our method outperforms the state-of-the-art methods on public VIPLoc datasets and the re-built CUC dataset. Xiao Wang 0029, Zheng Wang 0007, Wu Liu 0005, Xin Xu 0007, Qijun Zhao, Shin'ichi Satoh 0001 |
ACM Multimedia | 4 |
| 2022 | SAM: Self Attention Mechanism for Scene Text Recognition Based on Swin Transformer
Xiang Shuai, Xiao Wang 0029, Wei Wang 0170, Xin Yuan 0009, Xin Xu 0007 |
MMM (1) | 5 |
| 2022 | Cover: International Journal of Intelligent Systems, Volume 37 Issue 5 May 2022abstractCover Caption: The cover image is based on the Research Article Efficient virtual data search for annotationfree vehicle reidentification by Zhijing Wan et al., https://doi.org/10.1002/int.22829. Zhijing Wan, Xin Xu 0007, Zheng Wang 0007, Toshihiko Yamasaki, Xiaolong Zhang 0002, Ruimin Hu |
Int. J. Intell. Syst. | 2 |
| 2022 | Efficient virtual data search for annotation-free vehicle reidentificationabstractVehicle reidentification (re-ID) is the task of retrieving the same vehicle across nonoverlapping cameras, which has made significant progress with the help of abundant manually annotated real images. To avoid the time-consuming and tedious labeling of real images, virtual data sets with large-scale synthetic images have recently been constructed to perform annotation-free model training. However, current methods fail to exploit the potential of virtual data search, that is, searching valuable and representative virtual subdata set for efficient training. This paper presents a novel data sampling strategy from both semantic and feature levels to perform an effective data search. The semantic level determines the sample number of each vehicle identity via the consistency constraint of attribute distribution for source domain and target domain; while the feature level searches valuable and representative samples of each vehicle identity. To our knowledge, we are among the first attempts to search effective virtual data to perform annotation-free vehicle re-ID. Extensive cross-domain experiments from virtual vehicle re-ID data sets to real vehicle re-ID data sets show that our data sampling strategy can significantly reduce the training data volume and even boost the re-ID performance. Zhijing Wan, Xin Xu 0007, Zheng Wang 0007, Toshihiko Yamasaki, Xiaolong Zhang 0002, Ruimin Hu |
Int. J. Intell. Syst. | 2 |
| 2022 | BP Neural Network Retrieval for Remote Sensing Atmospheric Profile of Ground-Based Microwave RadiometerabstractVertical distributions of temperature and humidity are two essential factors for understanding the atmospheric structure, extreme weather events, and regional and global climate. The ground-based microwave radiometer (MWR), which acts as a passive sensor and operates continuously under all weather conditions, has an irreplaceable role in measuring the vertical information of the temperature and water content in the atmosphere. In this letter, we proposed a four-layer back-propagation neural network (BPNN) method to retrieve temperature and relative humidity (RH) profiles from the bright temperature measured by the MWR. In contrast to the traditional BPNN, this method has greater advantages in dealing with the problems of overfitting, gradient disappearance, and gradient explosion in vertical atmospheric retrieval. By adding dropout layers, it can also help to describe the nonlinear relationships for RH profiles. Results showed that the performance of the four-layer BPNN method was better than the quadratic regression (QR, provided by MWR manufacturer) method under both cloud and cloud-free conditions. Compared with measurements of radiosonde data, root-mean-square error of temperature and RH, BPNN achieves 1.88 K and 19.30% under cloud conditions and 2.03 K and 15.10% under cloud-free conditions, respectively, whereas the corresponding values by using the QR method were only 3.07 K and 24.28% under cloud conditions and 4.14 K and 18.96% under cloud-free conditions, respectively. Temperature and RH profiles retrieval with high precision have increased the efficiency of the MWR observations and provided a data foundation for further atmospheric climate research. Xin Xu 0007, Shikuan Jin, Yingying Ma 0001, Boming Liu, Wei Gong 0004 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Towards generalizable person re-identification with a bi-stream generative model
Xin Xu 0007, Wei Liu 0183, Zheng Wang 0007, Ruimin Hu |
Pattern Recognit. | 1 |
| 2022 | Continuous and Unified Person Re-IdentificationabstractPerson re-identification (ReID) aims to match pedestrian images across disjoint cameras. Mainstream Re-ID tasks focus on training ReID models once using all the data, which become limited in some real-world scenarios where training data tends to arrive in stages. To match scenarios where training data is incrementally available, some works began to explore ReID task that can make efficient use of piecemeal new data. However, due to the limitations of the training and testing setups, these efforts are still preliminary explorations. In this paper, we explore a novel yet harder Continuous and Unified ReID (CUReID), which not only enables to continuously learn discrimination knowledge from data streams with style differences, but also to be uniformly evaluated discriminatory capability on all the data (seen and unseen). Furthermore, we propose a novel Generalized Feature Decoupled Learning (GFDL) framework for CUReID, which characterizes by introducing alternate training with extra images to solve the problem of optimization divergence between regularisation (learning new knowledge) and generalization (anti-forgetting old knowledge) tasks. In our newly proposed benchmark setup, GFDL achieves the state-of-the-art performance. Zhu Mao, Xiao Wang 0029, Xin Xu 0007, Zheng Wang 0007, Chia-Wen Lin |
IEEE Signal Process. Lett. | 3 |
| 2022 | Full Coverage Estimation of the PM Concentration Across China Based on an Adaptive Spatiotemporal ApproachabstractParticulate pollution threatens the ecological environment, air quality, and public health. Therefore, it has become an increasing concern for the public and governments in recent decades. In this study, a full coverage PM2.5(aerodynamic diameter of less than 2.5 microns) estimation strategy is proposed based on spatiotemporal machine learning approaches including the Convolutional Neural Network with Long Short-Term Memory (CNN-LSTM) and Random Forest (RF). The RF estimates PM2.5by considering the features of a single pixel, while the introduction of the CNN-LSTM (size of 7 × 7 × 4) assists in exploiting the spatiotemporal correlation of surrounding pixel features. Compared with linear models and empirical spatiotemporal weight methods, our CNN-LSTM+RF avoids the uncertainty and complexity owing to actual measurements of the surrounding sites. In addition, full coverage is achieved using both satellite data and reanalysis data. Results showed that, the Root Mean Squared Error (RMSE) and coefficient of determination (R2) of the CNN-LSTM+RF were 12.790 μg/m3and 0.910, respectively, in sample-based Cross-Validation (CV). From the perspective of the season, the best performance of the CNN-LSTM+RF was found in autumn (R2of 0.915) and the lowest was in summer (R2of 0.848). In the meantime, for the different regions of China, the CNN-LSTM+RF also showed stable performance. The proposed method can generate high-precision continuous PM2.5distribution maps that provide beneficial support for improving environmental and public health, and provide a reference for using deeper networks. Cunxing Lei, Xin Xu 0007, Yingying Ma 0001, Shikuan Jin, Boming Liu, Wei Gong 0004 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Retrieving the Vertical Distribution of PM2.5 Mass Concentration From Lidar Via a Random Forest ModelabstractThe vertical distribution of fine particles with a diameter$ < 2.5~\mu \text{m}$(PM2.5) plays an important role in understanding the transport of air pollution and in making decisions regarding the prevention and control of regional air pollution. However, the studies of the vertical distribution of PM2.5were limited by the lack of monitoring data obtained with vertical sampling strategies. The lidar system can obtain the aerosol profile, which provides the possibility to measure PM2.5profile. Here, the vertical distributions of PM2.5concentrations were investigated on the basis of lidar data from January 2014 to October 2015. Linear regression, improved linear regression, and random forest (RF) models were used to retrieve the PM2.5concentration profile from lidar data. The models were built based on the relationship among extinction coefficient (EC), temperature ($T$), relative humidity (RH), and surface PM2.5mass concentration. Comparison of the estimated and observed PM2.5showed that the RF model exhibited the best inversion effect. The correlation coefficient reached 0.75, and the root mean absolute error (RMAE) and root mean square error (RMSE) were 3.94 and 21.1$\mu \text{g}/\text{m}^{3}$, respectively. Error analysis indicated that the estimated PM2.5retrieved using the linear and improved linear models (ILMs) was smaller than the observed PM2.5when EC was less than 0.7 km−1, whereas PM2.5was evidently overestimated during winter pollution days. The reason might be that the effects of$T$and RH were inaccurately considered. Finally, the seasonal variation of the PM2.5profiles was investigated. Results indicated that the mass concentration of PM2.5was relatively large within 0.5–1.5 km, with a maximum of 60$\mu \text{g}/\text{m}^{3}$. The findings obtained here provide guidance for PM2.5vertical observation and regional pollutant transport. Yingying Ma 0001, Boming Liu, Xin Xu 0007, Shikuan Jin, Wei Gong 0004 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Sampling and Re-Weighting: Towards Diverse Frame Aware Unsupervised Video Person Re-IdentificationabstractVideo person re-identification (re-ID) methods extract richer features from video tracklets than image-based ones and have received growing attention. However, existing supervised methods require numerous cross-camera identity labels, which is impractical for large-scale data. Although clustering-based unsupervised methods have been exploited to obtain pseudo labels and train the models iteratively for video person re-ID, they remain in their infancy due to the diversity of person images and uncertainty in the image quality of video tracklets. In this work, we employ two strategies ofSampling andRe-weighting forClustering (SRC) to obtain robust and discriminative person feature representations. This method considers the influence of two kinds of frames in the tracklet: 1) Detection errors or heavy occlusions generate noisy frames in the tracklet. These tracklets with noisy frames may be assigned with unreliable annotations during clustering. 2) Different frames are identified by the model with varying degrees of difficulty, caused by pose changes or partial occlusions. We call them hard frames, which are hard to identify but informative. To alleviate these problems, we propose a dynamic noise trimming module and diverse frame re-weighting module for sampling and re-weighting. The dynamic noise trimming module strengthens the dependability of the tracklet representation by removing noisy frames to enhance the clustering accuracy. The diverse frame re-weighting module focuses on training hard frames to enhance the learning of rich information from tracklet. Experiments on three video datasets,i.e.DukeMTMC-VideoReID, MARS and PRID2011, demonstrate the effectiveness of the proposed SRC under the unsupervised re-ID setting. Pengyu Xie, Xin Xu 0007, Zheng Wang 0007, Toshihiko Yamasaki |
IEEE Trans. Multim. | 2 |
| 2022 | Rank-in-Rank Loss for Person Re-identificationabstractPerson re-identification (re-ID) is commonly investigated as a ranking problem. However, the performance of existing re-ID models drops dramatically, when they encounter extreme positive-negative class imbalance (e.g., very small ratio of positive and negative samples) during training. To alleviate this problem, this article designs a rank-in-rank loss to optimize the distribution of feature embeddings. Specifically, we propose a Differentiable Retrieval-Sort Loss (DRSL) to optimize the re-ID model by ranking each positive sample ahead of the negative samples according to the distance and sorting the positive samples according to the angle (e.g., similarity score). The key idea of the proposed DRSL lies in minimizing the distance between samples of the same category along with the angle between them. Considering that the ranking and sorting operations are non-differentiable and non-convex, the DRSL also performs the optimization of automatic derivation and backpropagation. In addition, the analysis of the proposed DRSL is provided to illustrate that the DRSL not only maintains the inter-class distance distribution but also preserves the intra-class similarity structure in terms of angle constraints. Extensive experimental results indicate that the proposed DRSL can improve the performance of the state-of-the-art re-ID models, thus demonstrating its effectiveness and superiority in the re-ID task. Xin Xu 0007, Xin Yuan 0009, Zheng Wang 0007, Kai Zhang 0002, Ruimin Hu |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2021 | M2M: Learning to Enhance Low-Light Image from Model to Mobile FPGA
Wei Wang 0170, Wei Hu 0001, Xin Xu 0007 |
CGI | 4 |
| 2021 | Unsupervised Video Person Re-Identification via Noise and Hard Frame Aware ClusteringabstractUnsupervised video-based person re-identification (re-ID) methods extract richer features from video tracklets than image-based ones. The state-of-the-art methods utilize clustering to obtain pseudo-labels and train the models iteratively. However, they underestimate the influence of two kinds of frames in the tracklet: 1) noise frames caused by detection errors or heavy occlusions exist in the tracklet, which may be allocated with unreliable labels during clustering; 2) the tracklet also contains hard frames caused by pose changes or partial occlusions, which are difficult to distinguish but informative. This paper proposes a Noise and Hard frame Aware Clustering (NHAC) method. NHAC consists of a graph trimming module and a node re-sampling module. The graph trimming module obtains stable graphs by removing noise frame nodes to improve the clustering accuracy. The node re-sampling module enhances the training of hard frame nodes to learn rich tracklet information. Experiments conducted on two video-based datasets demonstrate the effectiveness of the proposed NHAC under the unsupervised re-ID setting. Pengyu Xie, Xin Xu 0007, Zheng Wang 0007, Toshihiko Yamasaki |
ICME | 2 |
| 2021 | Consistency-Constancy Bi-Knowledge Learning for Pedestrian Detection in Night SurveillanceabstractPedestrian detection in the night surveillance is a challenging yet not largely explored task. As the success of the detector in the daytime surveillance and the convenient acquisition of all-weather data, we learn knowledge from these data to benefit pedestrian detection in night surveillance. We find two key properties of surveillance: distribution cross-time consistency and background cross-frame constancy. This paper proposes a consistency-constancy bi-knowledge learning (CCBL) for pedestrian detection in night surveillance, which is able to simultaneously achieve the night pedestrian detection's useful knowledge, coming from day and night surveillance. Firstly, based on the robustness of the existing detector in day surveillance, we obtain pedestrians' distribution in the daytime scene using the detector's detection results in the daytime scene. Based on the consistency of pedestrians' distribution during the day and night in the same scene, the pedestrian distribution from daytime is used as the consistency-knowledge for pedestrian detection in night surveillance. Secondly, the background as a constant knowledge of the surveillance scene is extractable and contributes to the division of the foreground, which contains most of the pedestrian regions and helps in pedestrian detection for night surveillance. Finally, we add bi-knowledge representation to promote each other and merge them together as the final pedestrian representation. Through extensive experiments, our CCBL significantly outperforms the state-of-the-art methods on public pedestrian detection datasets. In the NightSurveillance dataset, CCBL reduced the average missed detection rate by 3.04% compared to the existing best method. Xiao Wang 0029, Zheng Wang 0007, Wu Liu 0005, Xin Xu 0007, Jing Chen 0003, Chia-Wen Lin |
ACM Multimedia | 4 |
| 2021 | Visible-Infrared Cross-Modal Person Re-identification based on Positive FeedbackabstractVisible-infrared person re-identification (VI-ReID) is undoubtedly a challenging cross-modality person retrieval task with increasing appreciation. Compared to traditional person ReID that focuses on person images in a single RGB mode, VI-ReID suffers from additional cross-modality discrepancy due to the different imaging processes of spectrum cameras. Several effective attempts have been made in recent years to narrow cross-modality gap aiming to improve the re-identification performance, but rarely study the key problem of optimizing the search results combined with relevant feedback. In this paper, we present the idea of cross-modality visible-infrared person re-identification combined with human positive feedback. This method allows the user to quickly optimize the search performance by selecting strong positive samples during the re-identification process. We have validated the effectiveness of our method on a public dataset, SYSU-MM01, and results confirmed that the proposed method achieved superior performance compared to the current state-of-the-art methods. Lingyi Lu, Xin Xu 0007 |
MMAsia | 2 |
| 2021 | Rethinking data collection for person re-identification: active redundancy reduction
Xin Xu 0007, Xiaolong Zhang 0002, Weili Guan, Ruimin Hu |
Pattern Recognit. | 1 |
| 2021 | Exploring Image Enhancement for Salient Object Detection in Low Light ImagesabstractLow light images captured in a non-uniform illumination environment usually are degraded with the scene depth and the corresponding environment lights. This degradation results in severe object information loss in the degraded image modality, which makes the salient object detection more challenging due to low contrast property and artificial light influence. However, existing salient object detection models are developed based on the assumption that the images are captured under a sufficient brightness environment, which is impractical in real-world scenarios. In this work, we propose an image enhancement approach to facilitate the salient object detection in low light images. The proposed model directly embeds the physical lighting model into the deep neural network to describe the degradation of low light images, in which the environment light is treated as a point-wise variate and changes with local content. Moreover, a Non-Local-Block Layer is utilized to capture the difference of local content of an object against its local neighborhood favoring regions. To quantitative evaluation, we construct a low light Images dataset with pixel-level human-labeled ground-truth annotations and report promising results on four public datasets and our benchmark dataset. Xin Xu 0007, Shiqin Wang, Zheng Wang 0007, Xiaolong Zhang 0002, Ruimin Hu |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2021 | Salient object detection from low contrast images based on local contrast enhancing and non-local feature learning
Tengda Guo, Xin Xu 0007 |
Vis. Comput. | 2 |
| 2021 | Deep Camera-Aware Metric Learning for Person ReidentificationabstractPerson reidentification (re‐id) suffers from a challenging issue due to the significant inconsistency of the camera network, including position, view, and brands. In this paper, we propose a deep camera‐aware metric learning (DCAML) model, where images on the identity‐level spaces are further projected into different camera‐level subspaces, which can explore the inherent relationship between identity and camera. Furthermore, we exploit dynamic training strategy to jointly multiple metrics for identity‐camera relationship learning and thus consumedly elevating the retrieval accuracy. Extensive experiments on the three public datasets demonstrated that our method performs competitive results compared to the state‐of‐the‐art person re‐id methods. Wei Liu 0183, Zhiqiang Hao, Xin Xu 0007 |
Wirel. Commun. Mob. Comput. | 5 |
| 2020 | Multi-granularity and Multi-semantic Model for Person Re-identification in Variable IlluminationabstractPerson re-identification (Re-ID) with deep neural networks has made tremendous improvements recently. The existing person Re-ID methods can do well in viewpoint, occlusion, resolution, etc. However, they are short of robustness under varying lighting conditions. The variable illumination will result in the inconsistency of color, contrast, and SNR (signal-noise ratio), which will cause lots of difficulties to identify the right person. This paper presents a multi-granularity and multi-semantic Re-ID model combined with image enhancement method to minimize the impact of illumination variations and optimize the feature extraction. The Retinex-based image enhancement method is used to balance the variable illumination and enhance the contour information of images. Furthermore, we add multi-granularity and multi-semantic layers in the network to extract powerful feature representation. The proposed model is evaluated on the Market-1501, DukeMTMC-reID and CUHK03 datasets. Extensive experiments show that the new deep neural network model can extract more robustness features from the enhanced images, and verify the effectiveness of our method under changing illumination conditions. Xin Xu 0007 |
SMC | 2 |
| 2020 | A survey of CAPTCHA technologies to distinguish between human and computer
Xin Xu 0007, Bo Li 0002 |
Neurocomputing | 1 |
| 2020 | End-to-end video subtitle recognition via a deep Residual Neural Network
Hongyu Yan, Xin Xu 0007 |
Pattern Recognit. Lett. | 2 |
| 2020 | Efficient Classification of Hot Spots and Hub Protein Interfaces by Recursive Feature Elimination and Gradient BoostingabstractProteins are not isolated biological molecules, which have the specific three-dimensional structures and interact with other proteins to perform functions. A small number of residues (hot spots) in protein-protein interactions (PPIs) play the vital role in bioinformatics to influence and control of biological processes. This paper uses the boosting algorithm and gradient boosting algorithm based on two feature selection strategies to classify hot spots with three common datasets and two hub protein datasets. First, the correlation-based feature selection is used to remove the highly related features for improving accuracy of prediction. Then, the recursive feature elimination based on support vector machine (SVM-RFE) is adopted to select the optimal feature subset to improve the training performance. Finally, boosting and gradient boosting (G-boosting) methods are invoked to generate classification results. Gradient boosting is capable of obtaining an excellent model by reducing the loss function in the gradient direction to avoid overfitting. Five datasets from different protein databases are used to verify our models in the experiments. Experimental results show that our proposed classification models have the competitive performance compared with existing classification methods. Xiaoli Lin, Xiaolong Zhang 0002, Xin Xu 0007 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2020 | Low-Light Image Enhancement With Semi-Decoupled DecompositionabstractLow-light image enhancement is important for high-quality image display and other visual applications. However, it is a challenging task as the enhancement is expected to improve the visibility of an image while keeping its visual naturalness. Retinex-based methods have well been recognized as a representative technique for this task, but they still have the following limitations. First, due to less-effective image decomposition or strong imaging noise, various artifacts can still be brought into enhanced results. Second, although the priori information can be explored to partially solve the first issue, it requires to carefully model the priori by a regularization term and usually makes the optimization process complicated. In this paper, we address these issues by proposing a novel Retinex-based low-light image enhancement method, in which the Retinex image decomposition is achieved in an efficient semi-decoupled way. Specifically, the illumination layer I is gradually estimated only with the input image S based on the proposed Gaussian Total Variation model, while the reflectance layer R is jointly estimated by S and the intermediate I. In addition, the imaging noise can be simultaneously suppressed during the estimation of R. Experimental results on several public datasets demonstrate that our method produces images with both higher visibility and better visual quality, which outperforms the state-of-the-art low-light enhancement methods in terms of several objective and subjective evaluation metrics. Shijie Hao, Yanrong Guo, Xin Xu 0007, Meng Wang 0001 |
IEEE Trans. Multim. | 4 |
| 2019 | Segmentation of Lesion in Dermoscopy Images Using Dense-Residual Network with Adversarial LearningabstractIn the field of medical images, skin lesion segmentation in dermoscopic images is a challenging task due to the irregular and blurring edges of the lesion and the presence of various artifacts. With the successful application of generative antagonistic network (GAN), a new neural network for skin lesion segmentation is proposed. The encoder-decoder with Dense-Residual block is used in the segmentation network which enables the network to be trained more efficiently. A multi-scale objective loss function is introduced to utilize deep supervision. We combine Jaccard distance and End Point Error which can solve lesion-background imbalance problem in pixel-level classification for skin lesion segmentation and also alleviate the problem of boundary ambiguity. A joint loss function is finally used, which includes a multi-scale objective loss function, End Point Error and Jaccard distance content loss function. Experiment results show that our algorithm is superior to other state-of-the-art algorithms on the ISBI2017. Wenli Tu, Xiaoming Liu 0004, Wei Hu 0001, Zhifang Pan, Xin Xu 0007, Bo Li 0002 |
ICIP | 5 |
| 2019 | Saliency Computational Model for Foggy Images by Fusing Frequency and Spatial CuesabstractA key challenge of saliency computation in foggy images is how to effectively detect salient objects which are less visible. The primary cause may lie in the fact that the light scattering through fog particles reduces image contrast. In this paper, we propose a frequency-spatial fusion saliency computational model based on discrete stationary wavelet transform (DSWT). The input image is firstly transformed into HSV color space, and the amplitude spectrum of each color channel is adjusted to generate the frequency domain saliency map. Then, the local-global superpixel contrast is measured to obtain the spatial domain saliency map. The DSWT is finally utilized to fuse the frequency-spatial cues. Experimental results indicate that the proposed model can efficiently reduce the influence of light scattering through fog particles, and can achieve the best performance in foggy images comparing to 16 state-of-the-art saliency models. Xin Xu 0007, Nan Mu, Li Chen 0011, Jing Tian 0002 |
ICIP | 1 |
| 2019 | BiRA-Net: Bilinear Attention Net for Diabetic Retinopathy GradingabstractDiabetic retinopathy (DR) is a common retinal disease that leads to blindness. For diagnosis purposes, DR image grading aims to provide automatic DR grade classification, which is not addressed in conventional research methods of binary DR image classification. Small objects in the eye images, like lesions and microaneurysms, are essential to DR grading in medical imaging, but they could easily be influenced by other objects. To address these challenges, we propose a new deep learning architecture, called BiRA-Net, which combines the attention model for feature extraction and bilinear model for fine-grained classification. Furthermore, in considering the distance between different grades of different DR categories, we propose a new loss function, called grading loss, which leads to improved training convergence of the proposed approach. Experimental results are provided to demonstrate the superior performance of the proposed approach. Ziyuan Zhao, Kerui Zhang, Xuejie Hao, Jing Tian 0002, Matthew Chua 0001, Li Chen 0011, Xin Xu 0007 |
ICIP | 7 |
| 2019 | Semisupervised Cross-Media Retrieval by Distance-Preserving Correlation Learning and Multi-modal Manifold Regularization
Hong Zhang 0022, Bo Li 0002, Xin Xu 0007 |
PRICAI (1) | 4 |
| 2019 | A spatial-frequency-temporal domain based saliency model for low contrast video sequences
Nan Mu, Xin Xu 0007, Xiaolong Zhang 0002 |
J. Vis. Commun. Image Represent. | 2 |
| 2019 | Finding autofocus region in low contrast surveillance images using CNN-based saliency algorithm
Nan Mu, Xin Xu 0007, Xiaolong Zhang 0002 |
Pattern Recognit. Lett. | 2 |
| 2019 | Efficiently Predicting Hot Spots in PPIs by Combining Random Forest and Synthetic Minority Over-Sampling TechniqueabstractHot spot residues bring into play the vital function in bioinformatics to find new medications such as drug design. However, current datasets are predominately composed of non-hot spots with merely a tiny percentage of hot spots. Conventional hot spots prediction methods may face great challenges towards the problem of imbalance training samples. This paper presents a classification method combining with random forest classification and oversampling strategy to improve the training performance. A strategy with an oversampling ability is used to generate hot spots data to balance the given training set. Random forest classification is then invoked to generate a set of forest trees for this oversampled training set. The final prediction performance can be computed recursively after the oversampling and training process. This proposed method is capable of randomly selecting features and constructing a robust random forest to avoid overfitting the training set. Experimental results from three data sets indicate that the performance of hot spots prediction has been significantly improved compared with existing classification methods. Xiaolong Zhang 0002, Xiaoli Lin, Jiafu Zhao, Xin Xu 0007 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2019 | Cross-Modality Retrieval by Joint Correlation LearningabstractAs an indispensable process of cross-media analyzing, comprehending heterogeneous data faces challenges in the fields of visual question answering (VQA), visual captioning, and cross-modality retrieval. Bridging the semantic gap between the two modalities is still difficult. In this article, to address the problem in cross-modality retrieval, we propose a cross-modal learning model with joint correlative calculation learning. First, an auto-encoder is used to embed the visual features by minimizing the error of feature reconstruction and a multi-layer perceptron (MLP) is utilized to model the textual features embedding. Then we design a joint loss function to optimize both the intra- and the inter-correlations among the image-sentence pairs, i.e., the reconstruction loss of visual features, the relevant similarity loss of paired samples, and the triplet relation loss between positive and negative examples. In the proposed method, we optimize the joint loss based on a batch score matrix and utilize all mutual mismatched paired samples to enhance its performance. Our experiments in the retrieval tasks demonstrate the effectiveness of the proposed method. It achieves comparable performance to the state-of-the-art on three benchmarks, i.e., Flickr8k, Flickr30k, and MS-COCO. Shuo Wang 0008, Dan Guo 0001, Xin Xu 0007, Li Zhuo 0001, Meng Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2018 | CBR Based Educational Method for the Postgraduate Course Image Processing and Machine Vision
Xin Xu 0007 |
ICIC (3) | 1 |
| 2018 | Screen-rendered text images recognition using a deep residual network based segmentation-free methodabstractText images recognition has long been known as a research hotspot of computer vision. However, screen-rendered text image pose great challenges to current character or text recognition methods due to its low resolution and low signal-to-noise ratio properties. In this paper, a segmentation-free method utilizing Residual Network (ResNet) and Recurrent Neural Network (RNN)-Connectionist Temporal Classification (CTC) is proposed to recognize Chinese and English texts in screen-rendered images. Text lines are firstly extracted from screen-rendered images to obtain feature sequences. Then, a bidirectional RNN layer is applied to model the contextual information within feature sequences and predict identification results. Finally, a CTC method is employed to calculate loss and yield the results. The proposed method can achieve the best performance on ORAND-CAR-A dataset, ORAND-CAR-B dataset and a generated dataset with the recognition accuracy of 91.89%, 93.79% and 95.67%, respectively. Moreover, experiments on several real screen-rendered text images also demonstrate the effectiveness of the proposed method. Xin Xu 0007, Hong Zhang 0022 |
ICPR | 1 |
| 2018 | Task scheduling with fault-tolerance in real-time heterogeneous systems
Jing Liu 0032, Mengxue Wei, Wei Hu 0001, Xin Xu 0007, Aijia Ouyang |
J. Syst. Archit. | 4 |
| 2018 | Cross-media retrieval based on semi-supervised regularization and correlation learning
Hong Zhang 0022, Du Tang, Xin Xu 0007 |
Multim. Tools Appl. | 4 |
| 2018 | Latent semantic factorization for multimedia representation learning
Hong Zhang 0022, Xin Xu 0007, Chunhua Deng |
Multim. Tools Appl. | 3 |
| 2018 | Discrete stationary wavelet transform based saliency information fusion from frequency and spatial domain in low contrast images
Nan Mu, Xin Xu 0007, Xiaolong Zhang 0002, Xiaoli Lin |
Pattern Recognit. Lett. | 2 |
| 2017 | Particle Swarm Optimization Based Salient Object Detection for Low Contrast Images
Nan Mu, Xin Xu 0007, Xiaolong Zhang 0002, Li Chen 0011 |
ICONIP (3) | 2 |
| 2016 | Covariance descriptor based convolution neural network for saliency computation in low contrast imagesabstractSaliency computational model with active environment perception can substantially facilitate a wide range of applications. Conventional saliency computational models primarily rely on hand-crafted low level image features, such as color or contrast. However, they may face great challenges in low lighting scenario, due to the lack of well-defined feature to represent saliency information in low contrast images. In this paper, we propose a novel deep neural network framework embedded with covariance descriptor for salient object detection in low contrast images. Several low-level features are extracted to compute their mutual covariance, which is then trained via a 7-layers convolutional neural network (CNN). The saliency map can be generated by estimating the final saliency score of each region via the pre-trained CNN model. Extensive experiments have been conducted on six challenging datasets to evaluate the performance of the proposed model against ten state-of-the-art models. Xin Xu 0007, Nan Mu, Xiaolong Zhang 0002, Bo Li 0002 |
IJCNN | 1 |
| 2016 | Hierarchical salient object detection model using contrast-based saliency and color spatial distribution
Xin Xu 0007, Nan Mu, Li Chen 0011, Xiaolong Zhang 0002 |
Multim. Tools Appl. | 1 |
| 2016 | Multiple kernel visual-auditory representation learning for retrieval
Hong Zhang 0022, Wenping Zhang, Wenhe Liu, Xin Xu 0007, Hehe Fan |
Multim. Tools Appl. | 4 |
| 2016 | A cross-media distance metric learning framework based on multi-view correlation mining and matching
Hong Zhang 0022, Xingyu Gao 0001, Xin Xu 0007 |
World Wide Web | 4 |
| 2015 | Hierarchical Features Fusion for Salient Object Detection in Low Contrast Images
Nan Mu, Xin Xu 0007 |
ICIC (2) | 2 |
| 2015 | Salient object detection from distinctive features in low contrast imagesabstractSaliency computational model with active environment perception can be useful for many applications including image segmentation, image compression, image retrieval, and etc. Conventional saliency computational models rely on handcrafted low level features, such as color or contrast. These models face great difficulties in low lighting scenarios, due to the lack of well-defined feature to interpret saliency information in low contrast images. In this paper, a new approach is proposed to detect salient object from low contrast images. The proposed approach explores the most distinguishable salient information in low contrast images based on low level features. Extensive experiments have been conducted to evaluate the performance of the proposed method against the state-of-the-art saliency computational models. Xin Xu 0007, Nan Mu, Hong Zhang 0022, Xiaowei Fu |
ICIP | 1 |
| 2015 | Nonnegative cross-media recoding of visual-auditory content for social media analysis
Hong Zhang 0022, Xin Xu 0007 |
Multim. Tools Appl. | 2 |
| 2014 | Block-Based Salient Region Detection Using a New Spatial-Spectral-Domain Contrast MeasureabstractVisual saliency is an important cue in human visual system, it can identify salient region in image. Image contrast has been utilized as an effective feature to detect the salient region. The conventional contrast measures utilize both spectral and spatial properties of image in many salient region detection methods. However, they only consider the local characteristics of image region, consequently, the global characteristics are neglected. This paper presented a new contrast measure by exploiting both local and global characteristics of image regions. Furthermore, the proposed measure is utilized to perform salient region detection in image. Experiments are conducted on the MSRA test database to compare the performance of the proposed approach with the state-of-the-art salient region detection algorithms. Nan Mu, Xin Xu 0007, Li Chen 0011, Jing Tian 0002 |
ISM | 2 |
| 2014 | A comparison of contrast measurements in passive autofocus systems for low contrast images
Xin Xu 0007, Xiaolong Zhang 0002, Shunxin Li, Xiaoming Liu 0004, Jinshan Tang |
Multim. Tools Appl. | 1 |
| 2012 | Mass Diagnosis in Mammography with Mutual Information Based Feature Selection and Support Vector Machine
Xiaoming Liu 0004, Bo Li 0002, Jun Liu 0011, Xin Xu 0007, Zhilin Feng |
ICIC (2) | 4 |
| 2012 | A Framework for GPS/INS based Portable Positioning SystemabstractIn this paper, we describe a framework for GPS/INS based Portable Positioning System. The framework includes two main components: receiving terminal and monitoring center. In the receiving terminal, the digital compass and the GPS modules are both connected with an I/O interface in a FPGA to provide an integrated positioning solution. This integrated positioning algorithm can solve the problems arising both in standalone modes and in traditional integrated measurements. In the monitoring center, several improvements have been discussed to ensure runtime efficiency and the robustness of the PPS. The experimental results from two positioning modes validate the effectiveness of the proposed algorithm. Xin Xu 0007, Heming Xu, Xiaoming Liu 0004, Jinshan Tang |
SMC | 1 |
| 2011 | Mass Classification with Level Set Segmentation and Shape Analysis for Breast Cancer Diagnosis Using Mammography
Xiaoming Liu 0004, Xin Xu 0007, Jun Liu 0011, Jinshan Tang |
ICIC (2) | 2 |
| 2011 | Adaptive Variance Based Sharpness Computation for Low Contrast Images
Xin Xu 0007, Jinshan Tang, Xiaolong Zhang 0002, Xiaoming Liu 0004 |
ICIC (1) | 1 |
| 2010 | Human behavior understanding for video surveillance: Recent advanceabstractWith the wide applications of video cameras in surveillance, video analysis technologies have attracted the attention from the researchers in computer vision field. In video analysis, human behavior recognition and understanding is an important research direction. By recognition and understanding the human behaviors, we can predict and recognize the happening of crimes and help to the police or other agencies to react immediately. In the past, large amount of intensive papers have been published on human behavior understanding in videos. Generally speaking, the procedure of human behavior understanding can be divided into the following stages: human segmentation and tracking, and human behavior recognition. In this paper, we provide a comprehensive survey of the recent development of all these stages. We will also discuss the difficulties in behavior understanding and identify possible future directions. Xin Xu 0007, Jinshan Tang, Xiaoming Liu 0004, Xiaolong Zhang 0002 |
SMC | 1 |
| 2010 | Web user behavior monitoring for campus networksabstractThe widespread networks in campus provide students with rich resources, but they also provide unhealthy information which may have negative effects on students. Filtering systems have been developed to prevent web users from unhealthy web contents. However, the false positive and false negative rates of those tools are still high and thus need to be improved. In this paper, we propose an Internet behavior monitoring method by integrating the BHO plug-in techniques with URL-based filtering technique, and apply this method in the campus network. Designed in the client/server mode, the method has the capabilities of filtering out unhealthy websites according to the given black list, and can record the students' online web activities into XML documents and/or pictures. The recorded contents can be replayed by the campus network administrators. This method assists teachers to keep eyes on students' learning state, and notifies the administrators to react immediately to help campus web users correct their inappropriate behaviors. Xiaolong Zhang 0002, Na Zeng, Jinshan Tang, Xin Xu 0007 |
SMC | 4 |