VLDB 2026 Research / reviewers in the wild / expert
Xianye Ben
dblp:96/10791
· DBLP profile ↗
43ranked-venue papers
12as first author
27since 2021 · last 2026
0000-0001-8083-3501ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 9 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 3 first-author · 12 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Security and privacy · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Lightweight Passive Depression Detection System Based on Interpretable Facial Feature AnalysisabstractAutomatic depression detection from facial videos is promising for IoT deployment but has long suffered from the black-box nature of deep models, which overlook how depression manifests on the face. This lack of interpretability leads to redundant feature learning, overly complex architectures, and consequently a trade-off between accuracy and deployability. We introduce a lightweight, interpretable system that explicitly models static–dynamic facial patterns, color channels, and regional textures. At its core is a Hybrid Static–Dynamic Model (HSDM) with detachable modules and identity decoupling, supported by two task-driven preprocessing steps (blue-channel suppression and 90° Gabor filtering). On AVEC2013/2014, our approach reduces mean absolute error by 10.1% on AVEC2014 while maintaining competitive results on AVEC2013. From a systems perspective, it achieves real-time inference on Jetson Nano (0.071M parameters, 0.21 GFLOPs, 17.69 ms latency, 56.54 samples/s throughput), outperforming larger SOTA models under identical conditions. The design facilitates on-device deployment and offers transparent decision paths through interpretable static–dynamic fusion, color synergy, and region-level analysis. Peng Zhang 0057, Tianhuan Huang, Jian Zhao 0002, Wei Xiang 0001, Xianye Ben |
IEEE Internet Things J. | 6 |
| 2026 | Cross-scale gated embedding graph model for skeleton-based anomaly detection
Peng Zhang 0057, Tianhuan Huang, Xianye Ben, Lei Chen 0095 |
Multim. Syst. | 4 |
| 2026 | Multi-view graph clustering via dual attention fusion and collaborative optimization
Naixuan Guo, Xuesheng Bian, Xiufang Xu, Shanliang Yao, Xianye Ben, Tian Zhou 0002 |
Neural Networks | 7 |
| 2026 | GaitADIB: Adversarial disentangled information bottleneck network for unseen-view gait recognition
Hanyue Du, Xianye Ben, Zunxiao Xu, Lei Chen 0095, Qiang Wu 0001 |
Pattern Recognit. | 2 |
| 2026 | High-Performance Accelerator for Constant-Time Cross-Domain Integer and Montgomery Inversion on FPGAabstractModular Inversion (MI) is one of the fundamental arithmetic operations in the finite field, which plays an essential role in various cryptographic applications and requires high performance and security. Unfortunately, the simple MI algorithm is vulnerable to side-channel attacks, such as the timing attack, which can compromise the cryptographic system by analyzing the time taken to execute cryptographic algorithms. Attackers may recover the initial data since the time can differ based on the input. Besides, the low complexity and low resource consumption of hardware implementations in MI are also challenging. In this article, we propose two novel modular inversion algorithms, named Constant-Time Integer Modular Inversion (CT-IMI) and Constant-Time Complementary Montgomery Modular Inversion (CT-CMMI). They both consist of constant iteration rounds to resist the timing attack. CT-IMI processes the data in the integer field, which is designed for common scenarios. CT-CMMI is suitable for the cross-domain case, which can directly use data in the Montgomery domain and avoid the conversion steps for some specific applications, e.g., scalar multiplication in Elliptic Curve Cryptography (ECC). In software simulations, we measure the average clock cycles for a single inversion and illustrate the relationship between various bit lengths and the latency. The significant differences between constant and non-constant algorithms demonstrate the vulnerability of modular inversion to timing attacks. In addition, we design two efficient hardware architectures on FPGA. Experimental results show that our CT-IMI can finish a single inversion in 2.56 \(\mu\) s with 4.2k LUTs, 1.8k FFs, and our CT-CMMI requires 2.45 \(\mu\) s with 2.7k LUTs, 1.6k FFs. The product of area and latency of our CT-IMI and CT-CMMI can reach 10.50 and 6.62, respectively, which shows optimal performance compared with all the results in the existing literature. Cheng Chen 0076, Gangqiang Yang, Hongchao Zhou, Hailiang Xiong, Xianye Ben, Zhiguo Wan |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2026 | SeeGait: Synergistic Co-Evolving Representations for Multimodal Gait Recognition via Hierarchical Multi-Stage FusionabstractGait recognition offers non-contact, long-distance identification but struggles with robustness against covariates like clothing variations, carrying conditions, and viewpoint changes. Existing methods predominantly rely on single modalities (e.g., silhouettes or skeletons) or employ shallow multimodal fusion, such as simple concatenation, which treats modalities as independent and static, failing to exploit their complementary strengths, shape cues from silhouettes and structural kinematics from skele-tons. To address these limitations, we introduce the Synergistic co-evolving representations (See) principle, enabling modalities to iteratively interact, guide, and refine each other across semantic hierarchies, fostering a unified, robust identity representation resilient to complex environments. This is realized through SeeGait, a novel multimodal framework featuring hierarchical multi-stage fusion. At its core, the Bidirectional Hierarchical Cross-Attention Synergy Module (BiHCASM) employs adaptive cross-modal attention to dynamically align and reweight features bidirectionally, allowing structural insights to enhance appearance focus and vice versa. Complementing this, the Hierarchical Spatiotemporal Transformer Encoder (HSTE) captures long-range skeleton dynamics, overcoming GCN limitations, while the Hierarchical Convolutional Silhouette Encoder (HCSE) extracts multi-scale silhouette pyramids for rich shape priors. Finally, a Holistic Feature Aggregation (HFA) strategy consolidates features from all stages for deep supervision, ensuring comprehensive optimization. By promoting mutual refinement, SeeGait mitigates covariate disruptions through enhanced complementarity, yielding superior discriminability. Extensive experiments show state-of-the-art performance, with 97.1% average Rank-1 accuracy on CASIA-B, and top results on CCPG and SUSTech1K. Hanyue Du, Xianye Ben, Xiankai Lu, Zunxiao Xu, Qingshuo Gao, Qiang Wu 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Human Identification at a Distance: Challenges, Methods and Results on the Competition HID 2025abstractHuman identification at a distance (HID) faces challenges due to the difficulty of acquiring traditional biometric modalities like face and fingerprints. Gait recognition offers a viable solution since it can be captured at a distance. To promote progress in gait recognition and provide a fair evaluation platform, the International Competition on Human Identification at a Distance (HID) has been organized annually since 2020. Since 2023, the competition has adopted the challenging SUSTech-Competition dataset, which includes significant variations in clothing, carried objects, and view angles. No training data is provided, requiring participants to train their models using external datasets. Each year, the competition applies a different random seed to generate distinct evaluation splits, reducing the risk of overfitting and ensuring fair evaluation of cross-domain generalization. Although the previous two competitions (HID 2023 and HID 2024) already utilized this dataset, HID 2025 aimed explicitly to explore whether algorithmic improvements could surpass the accuracy limits observed previously. Despite these heightened challenges, participants again demonstrated significant advancements, with the highest accuracy reaching 94.2%, setting a new benchmark for this dataset. We also analyze key technical trends and outline potential directions for future research on gait recognition. Jingzhe Ma, Jianlong Yu, Zunxiao Xu, Xue Cheng, Zepeng Wang 0002, Kazuki Osamura, Rujie Liu, Narishige Abe, Shunli Zhang 0005, Haojun Xie, Weiming Wu, Wenxiong Kang, Qingshuo Gao, Jiaming Xiong, Xianye Ben, Lei Chen 0095, Lichen Song, Junjian Cui, Haijun Xiong, Junhao Lu, Bin Feng 0001, Baoquan Zhao, Ke Xu 0001, Yongzhen Huang, Liang Wang 0001, Manuel J. Marín-Jiménez, Md. Atiqur Rahman Ahad, Shiqi Yu 0001 |
IJCB | 22 |
| 2025 | Mamba-SF: Monocular Scene Flow Learning with State Space ModelsabstractMonocular scene flow estimation has been a long-standing problem in computer vision. Methods based on the RAFT architecture are currently the mainstream approaches, while often overlooking the long-range dependencies in motion and texture features and fail to fully utilize the spatial information in texture features. In this paper, we consider that using Transformers introduces high computational complexity. Therefore, we propose the Mamba Motion Module based on State Space Models design, which first models long-range dependencies in motion and texture features and then fully leverages the spatial information in texture features to enhance the motion features, generating global motion features while maintaining low computational complexity. Additionally, texture features play a significant role in constraining motion boundaries. Therefore, we propose the Enhanced Texture Module, which predicts a set of channel weights from the global motion features to enrich the channel properties of the texture features and concatenates texture features with the global motion features along the channels, thereby improving scene flow accuracy. Experimental results show that our method achieves highly competitive results on the KITTI 2015 and Eigen Split datasets, increasing by 18.82% and 2.15% compared to the baseline, respectively. Xuezhi Xiang, Xianye Ben, Insha Hassan, Mingliang Zhai, Lei Zhang 0093, Xiantong Zhen |
ICIP | 3 |
| 2025 | Triplet Contrastive Learning with Learnable Sequence Augmentation for Sequential RecommendationabstractThe quality of augmented data directly affects the performance of contrastive learning. Low-quality augmentation offers limited benefits for model optimization. Existing contrastive learning-based sequential recommendation works primarily utilize heuristic data augmentation methods, which often exhibit excessive randomness and struggle to generate positive samples that align with users' true intentions. Wei Wang 0375, Yujie Lin 0001, Moyan Zhang, Jianli Zhao 0002, Xianye Ben, Pengjie Ren |
SIGIR | 7 |
| 2025 | EST transformer: enhanced spatiotemporal representation learning for time series anomaly detection
Xianye Ben, Lei Chen 0095 |
J. Intell. Inf. Syst. | 3 |
| 2025 | MSFIQA: A multi-scale feature fusion network based on human visual perception for no-reference image quality assessment
Tianfeng Xia, Yongcan Zhao, Peng Zhang 0057, Xianye Ben, Lei Chen 0095 |
J. Vis. Commun. Image Represent. | 4 |
| 2025 | Dual complementarity transformer for micro-expression recognition
Lei Chen 0095, Tianhuan Huang, Xianye Ben |
Multim. Syst. | 5 |
| 2025 | ALAD: A New Unsupervised Time Series Anomaly Detection Paradigm Based on Activation LearningabstractTime series anomaly detection has been received growing interest in industrial and academic communities due to its substantial theoretical value and practical significance in reality. Recent advanced methods for time series anomaly detection are based on deep learning techniques, since they have shown their superiority in some specific situations. However, most existing deep learning-based anomaly detection methods require predefined, specific tasks of reconstruction or prediction, necessitating task-specific loss functions. Designing such anomaly-aware loss functions poses a significant challenge due to the ambiguity in defining ground-truth anomalies. Moreover, these methods often rely on complex network architectures that tend to lead to over-generalization, resulting in even abnormal data being well reconstructed or fitted. To mitigate this situation, grounded in activation learning theory, we propose a novel unsupervised time series anomaly detection paradigm termed ALAD. ALAD utilizes a straightforward fully connected network architecture, measuring the typicality of input patterns through the sum of the squared output. Despite its simplicity, ALAD achieves competitive performance compared to state-of-the-art models trained using backpropagation. By utilizing various real-world and synthetic datasets, experimental results have confirmed the effectiveness and feasibility of the proposed paradigm. This work also demonstrates that biologically-plausible local learning can sometimes outperform backpropagation in real-world scenarios. Fengqian Ding, Xianye Ben, Hongchao Zhou |
IEEE Trans. Big Data | 3 |
| 2025 | Rethinking Appearance-Based Deep Gait Recognition: Reviews, Analysis, and Insights From Gait Recognition EvolutionabstractGait recognition is a prominent biometric recognition technique extensively employed in public security. Appearance-based and model-based gait recognition are two categories of methods commonly used. Specifically, appearance-based methods, which use silhouettes to represent body information, typically outperform model-based methods that rely on skeleton data, making them more popular. Recently, the shift from single-frame templates to multiframe silhouettes has advanced appearance-based gait recognition with better spatiotemporal representation. However, there is a notable lack of comprehensive studies that deepen the understanding of multiframe appearance-based gait recognition methods. This article reviews various methods to trace the evolution of gait recognition. Furthermore, we unify various performant models in one framework, study the overlooked effects on data arrangement, and explore the scaling ability of existing methods. Besides the advancement in gait recognition, we also summarize the current challenges and future prospects to foster future research. Changxin Ye, Wenzheng Xu, Xianye Ben, Fei-Yue Wang 0001, Junping Zhang |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | Privacy-Preserving Sequential Recommendation with Collaborative ConfusionabstractSequential recommendation has attracted a lot of attention from both academia and industry, however the privacy risks associated with gathering and transferring users’ personal interaction data are often underestimated or ignored. Existing privacy-preserving studies are mainly applied to traditional collaborative filtering or matrix factorization rather than sequential recommendation. Moreover, these studies are mostly based on differential privacy or federated learning, which often lead to significant performance degradation, or have high requirements for communication. In this work, we address privacy-preserving from a different perspective. Unlike existing research, we capture collaborative signals of neighbor interaction sequences and directly inject indistinguishable items into the target sequence before the recommendation process begins, thereby increasing the perplexity of the target sequence. Even if the target interaction sequence is obtained by attackers, it is difficult to discern which ones are the actual user interaction records. To achieve this goal, we introduce a novel sequential recommender system called CoLlaborative-cOnfusion seqUential recommenDer (CLOUD) , which incorporates a collaborative confusion mechanism to modify the raw interaction sequences before conducting recommendation. Specifically, CLOUD first calculates the similarity between the target interaction sequence and other neighbor sequences to find similar sequences. Then, CLOUD considers the shared representation of the target sequence and similar sequences to determine the operation to be performed: keep, delete, or insert. A copy mechanism is designed to make items from similar sequences have a higher probability to be inserted into the target sequence. Finally, the modified sequence is used to train the recommender and predict the next item. We conduct extensive experiments on three benchmark datasets. The experimental results show that CLOUD achieves a maximum modification rate of 66.57% on interaction sequences and obtains over 99% recommendation accuracy compared to the state-of-the-art sequential recommendation methods. This proves that CLOUD can effectively protect user privacy at minimal recommendation performance cost, which provides a new solution for privacy-preserving for sequential recommendation. Our implementation is available at https://github.com/weiwang0927/CLOUD . Wei Wang 0375, Yujie Lin 0001, Pengjie Ren, Zhumin Chen, Tsunenori Mine, Jianli Zhao 0002, Qiang Zhao 0011, Moyan Zhang, Xianye Ben |
ACM Trans. Inf. Syst. | 9 |
| 2024 | Efficient abnormal behavior detection with adaptive weight distribution
Yefeng Qin, Lei Chen 0095, Peng Zhang 0057, Xianye Ben |
Neurocomputing | 5 |
| 2024 | You watch once more: a more effective CNN architecture for video spatio-temporal action localization
Yefeng Qin, Lei Chen 0095, Xianye Ben |
Multim. Syst. | 3 |
| 2024 | Boosting Micro-Expression Recognition via Self-Expression Reconstruction and Memory Contrastive LearningabstractMicro-expression (ME) is an instinctive reaction that is not controlled by thoughts. It reveals one's inner feelings, which is significant in sentiment analysis and lie detection. Since micro-expression is expressed as subtle facial changes within particular facial action units, learning discriminative and generalized features for Micro-expression Recognition (MER) is challenging. To achieve the purpose, this paper proposes a novel MER framework that simultaneously integrates supervised Prototype-based Memory Contrastive Learning (PMCL) for discriminative feature mining and adds Self-expression Reconstruction (SER) as an auxiliary task and regularization for better generalization. In particular, the proposed SER module is forced as a regularization by reconstructing input ME from the randomly dropped patch- wise features in the bottleneck. And, the PMCL module globally compares historical and current cluster agents learned from training instances to enhance intra-class compactness and inter-class separability. Extensive experiments are conducted on three benchmarks, e.g., SMIC, CASME II, and SAMM, under evaluation criteria of both Composite Database Evaluation (CDE) and Single Database Evaluation (SDE) protocols. The results show our method surpasses other state-of-the-art approaches under various evaluation metrics, achieving overall 86.30% unweighed F1-score and 88.30% unweighed average recall on the composite dataset. Furthermore, the ablation studies verify the effectiveness of our SER for better generalization and PMCL for better discrimination in learning feature representation from limited micro-expression samples. Yongtang Bao, Peng Zhang 0057, Caifeng Shan, Xianye Ben |
IEEE Trans. Affect. Comput. | 6 |
| 2024 | GaitDAN: Cross-View Gait Recognition via Adversarial Domain AdaptationabstractView change causes significant differences in the gait appearance. Consequently, recognizing gait in cross-view scenarios is highly challenging. Most recent approaches either convert the gait from the original view to the target view before recognition is carried out or extract the gait feature irrelevant to the camera view through either brute force learning or decouple learning. However, these approaches have many constraints, such as the difficulty of handling unknown camera views. This work treats the view-change issue as a domain-change issue and proposes to tackle this problem through adversarial domain adaptation. This way, gait information from different views is regarded as the data from different sub-domains. The proposed approach focuses on adapting the gait feature differences caused by such sub-domain change and, at the same time, maintaining sufficient discriminability across the different people. For this purpose, a Hierarchical Feature Aggregation (HFA) strategy is proposed for discriminative feature extraction. By incorporating HFA, the feature extractor can well aggregate the spatial-temporal feature across the various stages of the network and thereby comprehensive gait features can be obtained. Then, an Adversarial View-change Elimination (AVE) module equipped with a set of explicit models for recognizing the different gait viewpoints is proposed. Through the adversarial learning process, AVE would not be able to identify the gait viewpoint in the end, given the gait features generated by the feature extractor. That is, the adversarial domain adaptation mitigates the view change factor, and discriminative gait features that are compatible with all sub-domains are effectively extracted. Extensive experiments on three of the most popular public datasets, CASIA-B, OULP, and OUMVLP richly demonstrate the effectiveness of our approach. Tianhuan Huang, Xianye Ben, Chen Gong 0002, Wenzheng Xu, Qiang Wu 0001, Hongchao Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Micro-expression action unit recognition based on dynamic image and spatial pyramid
Guanqun Zhou, Shusen Yuan, Hongbo Xing, Youjun Jiang, Pinyong Geng, Yewen Cao, Xianye Ben |
J. Supercomput. | 7 |
| 2023 | Tackling Micro-Expression Data Shortage via Dataset Alignment and Active LearningabstractThe research on micro-expression recognition has been drawing great attention in recent years, because of its great potential in the lie detection, clinical diagnosis, and national security. Amongst many challenges, data shortage stands out as it directly prevents an accurate training of micro-expression recognition algorithm. In this work, we present our approach within a dataset alignment and active learning (DAAL) framework. DAAL effectively queries minimum examples to label, as well as transfers features from micro-expression dataset to macro-expression dataset. Specifically, the features from micro-expression dataset are mapped to the macro-expression dataset with a translator, so that the classifier trained in macro-expression dataset can be adjusted and adapted to boost the classification performance on the micro-expression dataset. Besides, the most informative examples in the micro-expression dataset are selected through active learning in an iterative way, which effectively improves the classification ability of the model. Comprehensive experiments on CASME, CASME II, SAMM and SMIC databases firmly demonstrate that the proposed DAAL outperforms previous works by a large margin on micro-expression recognition task. Xianye Ben, Chen Gong 0002, Tianhuan Huang, Chuanye Li |
IEEE Trans. Multim. | 1 |
| 2022 | Decomposing Identity and View for Cross-View Gait RecognitionabstractOne of many challenges in cross-view gait recognition is that a pedestrian's gaits have huge visual differences under different views. A promising solution to this challenge is to extract fundamental discriminative view-invariant gait features, so that gait recognition systems can be robust to viewpoint varying. In order to extract discriminative gait features, this paper proposes a method to decompose identity and view for cross-view gait recognition. Our method employs a newly designed auto-encoder to detach the identity features from the view features. We train our model with the classification and metric learning to obtain discriminative identity feature. Also, both the view regression and identity ambiguity are applied to ensure that the view features only contain view information. Finally, a dissimilarity loss is proposed to increase the distribution discrepancy between identity feature and view feature, so as to further refine the identity feature. The experimental results on the CASIA-B and OU-ISIR databases show that the proposed method performs well in cross-view gait recognition. Xinliang Zhai, Xianye Ben, Tingxuan Xie |
ICME | 2 |
| 2022 | A novel micro-expression detection algorithm based on BERT and 3DCNN
Yanxin Song, Lei Chen 0095, Yang Chen 0063, Xianye Ben, Yewen Cao |
Image Vis. Comput. | 5 |
| 2022 | Unsupervised cross-database micro-expression recognition based on distribution adaptation
Ruixue Xiao, Xianye Ben, Kidiyo Kpalma, Hongchao Zhou |
Multim. Syst. | 5 |
| 2022 | Video-Based Facial Micro-Expression Analysis: A Survey of Datasets, Features and AlgorithmsabstractUnlike the conventional facial expressions, micro-expressions are involuntary and transient facial expressions capable of revealing the genuine emotions that people attempt to hide. Therefore, they can provide important information in a broad range of applications such as lie detection, criminal detection, etc. Since micro-expressions are transient and of low intensity, however, their detection and recognition is difficult and relies heavily on expert experiences. Due to its intrinsic particularity and complexity, video-based micro-expression analysis is attractive but challenging, and has recently become an active area of research. Although there have been numerous developments in this area, thus far there has been no comprehensive survey that provides researchers with a systematic overview of these developments with a unified evaluation. Accordingly, in this survey paper, we first highlight the key differences between macro- and micro-expressions, then use these differences to guide our research survey of video-based micro-expression analysis in a cascaded structure, encompassing the neuropsychological basis, datasets, features, spotting algorithms, recognition algorithms, applications and evaluation of state-of-the-art approaches. For each aspect, the basic techniques, advanced developments and major challenges are addressed and discussed. Furthermore, after considering the limitations of existing micro-expression datasets, we present and release a new dataset — calledmicro-and-macro expression warehouse(MMEW) — containing more video samples and more labeled emotion types. We then perform a unified comparison of representative methods on CAS(ME)$^2$for spotting, and on MMEW and SAMM for recognition, respectively. Finally, some potential future research directions are explored and outlined. Xianye Ben, Junping Zhang, Kidiyo Kpalma, Weixiao Meng 0001, Yong-Jin Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Enhanced Spatial-Temporal Salience for Cross-View Gait RecognitionabstractGait recognition can be used in person identification and re-identification by itself or in conjunction with other biometrics. Although gait has both spatial and temporal attributes, and it has been observed that decoupling spatial feature and temporal feature can better exploit the gait feature on the fine-grained level. However, the spatial-temporal correlations of gait video signals are also lost in the decoupling process. Direct 3D convolution approaches can retain such correlations, but they also introduce unnecessary interferences. Instead of common 3D convolution solutions, this paper proposes an integration of decoupling process into a 3D convolution framework for cross-view gait recognition. In particular, a novel block consisting of a Parallel-insight Convolution layer integrated with a Spatial-Temporal Dual-Attention (STDA) unit is proposed as the basic block for global spatial-temporal information extraction. Under the guidance of the STDA unit, this block can well integrate spatial-temporal information extracted by two decoupled models and at the same time retain the spatial-temporal correlations. In addition, a Multi-Scale Salient Feature Extractor is proposed to further exploit the fine-grained features through context awareness extension of part-based features and adaptively aggregating the spatial features. Extensive experiments on three popular gait datasets, namely CASIA-B, OULP and OUMVLP, demonstrate that the proposed method outperforms state-of-the-art methods. Tianhuan Huang, Xianye Ben, Chen Gong 0002, Baochang Zhang 0001, Qiang Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Learning Spatial-Temporal Representations Over Walking Tracklet for Long-Term Person Re-Identification in the WildabstractLong-term person re-identification (re-ID) aims to build identity correspondence of the Target Subject of Interest (TSI) exposed under surveillance cameras over a long time interval. Compared to the conventional short-term re-ID studied by most existing works, it suffers an additional problem: significant dressing change observed with time lapsing. Unfortunately, this variation in long-term person re-ID case contradicts the assumption of prior short-term re-ID approaches, and thus causes significant difficulties if conventional short-term re-ID methods are applied. To address the problem, this paper proposes to learn hybrid feature representation via a two-stream network named SpTSkM, including a spatial-temporal stream and a skeleton motion stream. The former performs directly on image sequences, which tends to learn identity-related spatial-temporal patterns such as body geometric structure and body movement. The latter operates on normalized 3D skeletons by adapting graph convolutional network, which tends to learn pure motion patterns from skeleton sequences. Both streams extract fine-grained level time-gap stable information that is robust to appearance changes in long-term re-ID and meanwhile maintains sufficient discriminability to differentiate different people. The final matching metric is obtained by mixing information of the two streams in a score-level fusion strategy. In addition, we collect a Cloth-Varying vIDeo re-ID (CVID-reID) dataset particularly for long-term re-ID. It contains video tracklets of celebrities posted on the Internet. These videos are snapshots under extremely different scenarios that include highly dynamic background, diverse camera views and abundant cloth variations on each TSI. These factors cause CVID-reID more complicated and closer to practice. Our experiments demonstrate the difficulty of long-term person re-ID and also validate the effectiveness of the proposed SpTSkM, showing the best performance. Peng Zhang 0057, Jingsong Xu, Qiang Wu 0001, Yan Huang 0023, Xianye Ben |
IEEE Trans. Multim. | 5 |
| 2020 | DAVD-Net: Deep Audio-Aided Video Decompression of Talking HeadsabstractClose-up talking heads are among the most common and salient object in video contents, such as face-to-face conversations in social media, teleconferences, news broadcasting, talk shows, etc. Due to the high sensitivity of human visual system to faces, compression distortions in talking heads videos are highly visible and annoying. To address this problem, we present a novel deep convolutional neural network (DCNN) method for very low bit rate video reconstruction of talking heads. The key innovation is a new DCNN architecture that can exploit the audio-video correlations to repair compression defects in the face region. We further improve reconstruction quality by embedding into our DCNN the encoder information of the video compression standards and introducing a constraining projection module in the network. Extensive experiments demonstrate that the proposed DCNN method outperforms the existing state-of-the-art methods on videos of talking heads. Xi Zhang 0019, Xiaolin Wu 0001, Xinliang Zhai, Xianye Ben, Chengjie Tu |
CVPR | 4 |
| 2020 | Coupled Bilinear Discriminant Projection for Cross-View Gait RecognitionabstractA problem that hinders good performance of general gait recognition systems is that the appearance features of gaits are more affected-prone by views than identities, especially when the walking direction of the probe gait is different from the register gait. This problem cannot be solved by traditional projection learning methods because these methods can learn only one projection matrix, and thus for the same subject, it cannot transfer cross-view gait features into similar ones. This paper presents an innovative method to overcome this problem by aligning gait energy images (GEIs) across views with the coupled bilinear discriminant projection (CBDP). Specifically, the CBDP generates the aligned gait matrix features for two views with two sets of bilinear transformation matrices, so that the original GEIs' spatial structure information can be preserved. By iteratively maximizing the ratio of inter-class distance metric to intra-class distance metric, the CBDP can learn the optimal matrix subspace where the GEIs across views are aligned in both horizontal and vertical coordinates. Therefore, the CBDP is also able to avoid the under-sample problem. We also theoretically prove that the upper and lower bounds of the objective function sequence of the CBDP are both monotonically increasing, so the convergence of the CBDP is demonstrated. In the terms of accuracy, the comparative experiments on the CASIA (B) and OU-ISIR gait databases show that our method is superior to the state-of-the-art cross-view gait recognition methods. More impressively, encouraging performance is obtained by our method even in matching a lateral-view gait with a frontal-view gait. Xianye Ben, Chen Gong 0002, Peng Zhang 0057, Qiang Wu 0001, Weixiao Meng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | A novel construction method of convolutional neural network model based on data-driven
Guofeng Zou, Guixia Fu, Mingliang Gao 0001, Jin Shen, Liju Yin, Xianye Ben |
Multim. Tools Appl. | 6 |
| 2019 | A general tensor representation framework for cross-view gait recognition
Xianye Ben, Peng Zhang 0057, Zhihui Lai 0001, Xinliang Zhai, Weixiao Meng 0001 |
Pattern Recognit. | 1 |
| 2019 | Coupled Patch Alignment for Matching Cross-View GaitsabstractGait recognition has attracted growing attention in recent years as the gait of humans has a strong discriminative ability even under low resolution at a distance. Unfortunately, the performance of gait recognition can be largely affected by view change. To address this problem, we propose a Coupled Patch Alignment (CPA) algorithm that effectively matches a pair of gaits across different views. To realize CPA, we first build a certain amount of patches, and each of them is made up of a sample as well as its intra-class and inter-class nearest-neighbors. Then we design an objective function for each patch to balance the cross-view intra-class compactness and the cross-view inter-class separability. Finally, all the local independent patches are combined to render a unified objective function. Theoretically, we show that the proposed CPA has a close relationship with Canonical Correlation Analysis (CCA). Algorithmically, we extend CPA to "Multi-dimensional Patch Alignment" (MPA) that can handle an arbitrary number of views. Comprehensive experiments on CASIA(B), USF and OU-ISIR gait databases firmly demonstrate the effectiveness of our methods over other existing popular methods in terms of cross-view gait recognition. Xianye Ben, Chen Gong 0002, Peng Zhang 0057, Xitong Jia, Qiang Wu 0001, Weixiao Meng 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | Coupled source domain targetized with updating tag vectors for micro-expression recognition
Xuena Zhu, Xianye Ben, Shigang Liu, Weixiao Meng 0001 |
Multim. Tools Appl. | 2 |
| 2018 | Learning effective binary descriptors for micro-expression recognition transferred by macro-information
Xianye Ben, Xitong Jia, Weixiao Meng 0001 |
Pattern Recognit. Lett. | 1 |
| 2017 | Accurately predicting heat transfer performance of ground heat exchanger for ground-coupled heat pump systems using data mining methods
Zhaoyi Zhuang, Xianye Ben, Jianhua Pang |
Neural Comput. Appl. | 2 |
| 2016 | On the distance metric learning between cross-domain gaits
Xianye Ben, Peng Zhang 0057, Weixiao Meng 0001, Wenhe Liu |
Neurocomputing | 1 |
| 2016 | An adaptive neural networks formulation for the two-dimensional principal component analysis
Xianye Ben, Weixiao Meng 0001 |
Neural Comput. Appl. | 1 |
| 2016 | Gait recognition and micro-expression recognition based on maximum margin projection with tensor representation
Xianye Ben, Peng Zhang 0057, Guodong Ge |
Neural Comput. Appl. | 1 |
| 2016 | A novel label learning algorithm for face recognition
Shigang Liu, Xianye Ben, Wankou Yang, Guoyong Qiu |
Signal Process. | 3 |
| 2015 | Low-resolution degradation face recognition over long distance based on CCA
Wankou Yang, Xianye Ben |
Neural Comput. Appl. | 3 |
| 2013 | Kernel coupled distance metric learning for gait recognition and face recognition
Xianye Ben, Weixiao Meng 0001 |
Neurocomputing | 1 |
| 2012 | Dual-ellipse fitting approach for robust gait periodicity detection
Xianye Ben, Weixiao Meng 0001 |
Neurocomputing | 1 |
| 2012 | An improved biometrics technique based on metric learning approach
Xianye Ben, Weixiao Meng 0001 |
Neurocomputing | 1 |