VLDB 2026 Research / reviewers in the wild / expert
Chen Chen 0036
dblp:65/4423-36
· DBLP profile ↗
52ranked-venue papers
5as first author
35since 2021 · last 2026
0000-0001-8297-6549ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 4 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 32 · 5 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HGATSolver: A Heterogeneous Graph Attention Solver for Fluid-Structure InteractionabstractFluid–structure interaction (FSI) systems involve distinct physical domains, fluid and solid, governed by different partial differential equations and coupled at a dynamic interface. While learning-based solvers offer a promising alternative to costly numerical simulations, existing methods struggle to capture the heterogeneous dynamics of FSI within a unified framework. This challenge is further exacerbated by inconsistencies in response across domains due to interface coupling and by disparities in learning difficulty across fluid and solid regions, leading to instability during prediction. To address these challenges, we propose the Heterogeneous Graph Attention Solver (HGATSolver). HGATSolver encodes the system as a heterogeneous graph, embedding physical structure directly into the model via distinct node and edge types for fluid, solid, and interface regions. This enables specialized message-passing mechanisms tailored to each physical domain. To stabilize explicit time stepping, we introduce a novel physics-conditioned gating mechanism that serves as a learnable, adaptive relaxation factor. Furthermore, an Inter-domain Gradient-Balancing Loss dynamically balances the optimization objectives across domains based on predictive uncertainty. Extensive experiments on two constructed FSI benchmarks and a public dataset demonstrate that HGATSolver achieves state-of-the-art performance, establishing an effective framework for surrogate modeling of coupled multi-physics systems. Haichuan Lin, Linying Cao, Xiao-Hu Zhou, Chen Chen 0036, Shuang-Yi Wang, Zeng-Guang Hou |
AAAI | 7 |
| 2026 | Continuous Shape-to-Texture Face Aging With Flow-Based Prior Latent Age Modulation and Attentional Alignment StyleGANabstractAlbeit recent Generative Models have achieved notable progress in synthesizing realistic facial aging images, many of them, e.g., GAN-based methods, cannot accurately capture the continuous progression of age-related shape-to-texture changes over time. In this paper, we propose an innovative facial age transformation framework that enables the generation of continuous shape-to-texture aging facial images. Firstly, the Prior Latent Age Modulation (PLAM) is designed to leverage the advantages of continuous sampling in high-dimensional space by normalizing flows to achieve precise and reversible mapping between the age attribute variable distributions and the prior latent space, ensuring smooth transitions along with facial aging. Secondly, we introduce the Attentional Feature Fusion (AFF), which dynamically allocates weights to effectively fuse the age attribute features by the latent space manipulation with the content features in StyleGAN, thereby generating facial images that accurately depict facial characteristics from shape to texture corresponding to specific ages. Finally, through quantitative and qualitative analysis of existing datasets, we validate the effectiveness and superiority of our proposed method in facial aging tasks. Xiyuan Hu, Jinglei Qu, Chen Chen 0036, Yu Liang 0003 |
IEEE Trans. Image Process. | 3 |
| 2025 | SemStereo: Semantic-Constrained Stereo Matching Network for Remote SensingabstractSemantic segmentation and 3D reconstruction are two fundamental tasks in remote sensing, typically treated as separate or loosely coupled tasks. Despite attempts to integrate them into a unified network, the constraints between the two heterogeneous tasks are not explicitly modeled, since the pioneering studies either utilize a loosely coupled parallel structure or engage in only implicit interactions, failing to capture the inherent connections. In this work, we explore the connections between the two tasks and propose a new network that imposes semantic constraints on the stereo matching task, both implicitly and explicitly. Implicitly, we transform the traditional parallel structure to a new cascade structure termed Semantic-Guided Cascade structure, where the deep features enriched with semantic information are utilized for the computation of initial disparity maps, enhancing semantic guidance. Explicitly, we propose a Semantic Selective Refinement (SSR) module and a Left-Right Semantic Consistency (LRSC) module. The SSR refines the initial disparity map under the guidance of the semantic map. The LRSC ensures semantic consistency between two views via reducing the semantic divergence after transforming the semantic map from one view to the other using the disparity map. Experiments on the US3D and WHU datasets demonstrate that our method achieves state-of-the-art performance for both semantic segmentation and stereo matching. Chen Chen 0036, Liangjin Zhao, Yuanchun He, Yingxuan Long, Kaiqiang Chen, Zhirui Wang 0003, Yanfeng Hu, Xian Sun 0001 |
AAAI | 1 |
| 2025 | Masked Spectrum ViT with Cross-Scale Fusion for Universal Deepfake DetectionabstractWith the rapid evolution of deepfake technologies, existing universal detection methods often fail to generalize when faced with new forgery patterns, as they tend to overfit specific artifacts found in the training data. To address this challenge, we propose a detection framework that integrates adaptive spectrum masking and cross-scale feature fusion, termed Masked Spectrum Vision Transformer with Cross-scale Fusion (MSViT-CF). Our approach introduces two key innovations: (1) Multi-band Artifact Enhancement Module (MAEM) reconstructs input images through wavelet decomposition and strategically perturbs high-frequency subbands via adaptive masking, amplifying subtle forgery traces while forcing the model to learn generalized artifact representations; (2) Dynamic Scale Fusion Transformer (DSFT) integrates multi-resolution frequency features through parallel convolutional-transformer pathways, dynamically weighting local spectrum anomalies and global structural inconsistencies. MAEM enhances artifact sensitivity through frequency-space discrepancy learning, while DSFT establishes cross-scale relationships between pixel-level irregularities and semantic-level inconsistencies via learnable attention gates. Experimental results demonstrate that MSViT-CF significantly outperforms existing state-of-the-art methods in detecting deepfake images generated by various GANs and diffusion models, exhibiting superior universality and robustness. Kaiwen Xu, Xiyuan Hu, Chen Chen 0036 |
ECAI | 3 |
| 2025 | SA-Occ: Satellite-Assisted 3D Occupancy Prediction in Real WorldabstractExisting vision-based 3D occupancy prediction methods are inherently limited in accuracy due to their exclusive reliance on street-view imagery, neglecting the potential benefits of incorporating satellite views. We propose SA-Occ, the first Satellite-Assisted 3D occupancy prediction model, which leverages GPS & IMU to integrate historical yet readily available satellite imagery into real-time applications, effectively mitigating limitations of ego-vehicle perceptions, involving occlusions and degraded performance in distant regions. To address the core challenges of cross-view perception, we propose: 1) Dynamic-Decoupling Fusion, which resolves inconsistencies in dynamic regions caused by the temporal asynchrony between satellite and street views; 2) 3D-Proj Guidance, a module that enhances 3D feature extraction from inherently 2D satellite imagery; and 3) Uniform Sampling Alignment, which aligns the sampling density between street and satellite views. Evaluated on Occ3D-nuScenes, SA-Occ achieves state-of-the-art performance, especially among single-frame methods, with a 39.05% mIoU (a 6.97% improvement), while incurring only 6.93 ms of additional latency per frame. Our code and newly curated dataset are available at https://github.com/chenchen235/SA-Occ. Chen Chen 0036, Zhirui Wang 0003, Taowei Sheng, Yundu Li, Peirui Cheng, Luning Zhang, Kaiqiang Chen, Yanfeng Hu, Xue Yang 0005, Xian Sun 0001 |
ICCV | 1 |
| 2025 | Model-Free Catheter Delivery Strategy for Robotic Transcatheter Tricuspid Valve ReplacementabstractTranscatheter tricuspid valve replacement (TTVR) has emerged as a promising minimally invasive procedure for treating severe tricuspid regurgitation (TR). However, accurate catheter delivery remains a significant challenge, primarily due to the reliance on 2D vision feedback, complex catheter kinematics, camera-to-robot pose calibration, which are difficult to generalize across patients. To address these issues, this paper presents a model-free robotic catheter delivery strategy for TTVR using Data-Enabled Predictive Control (DeePC). This approach leverages data-driven control to optimize catheter positioning without the need for prior knowledge of the system’s dynamics, eliminating the need for complex kinematic models or camera calibration. The proposed method incorporates environmental constraints to ensure the safety of the procedure, delivering the catheter to the desired location with high accuracy across varying catheters and camera poses. Experimental results demonstrate the effectiveness and versatility of the approach, suggesting its potential for broader applications in robotic-assisted surgeries. This work presents a new perspective for vision based robotic TTVR, as well as other clinical interventions involving robotic catheter control. Haichuan Lin, Longyue Tan, Weizhao Wang, Yuen Chiu Ng, Xilong Hou 0001, Chen Chen 0036, Xiao-Hu Zhou, Zeng-Guang Hou, Shuangyi Wang |
IROS | 9 |
| 2025 | COD-SAM: Camouflage object detection using SAM
Dongyang Gao, Chen Chen 0036, Xiyuan Hu |
Pattern Recognit. | 4 |
| 2024 | Global-Local Unified Cross-View Enhancing Framework for Person Re-IdentificationabstractIn Pedestrian Re-identification (ReID) tasks, extracting robust and discriminative features is a key challenge. Recent studies have mainly concentrated on the extraction of features from single images. However, limited research has been done on considering the relationships between different images, which is crucial in the ReID task. To effectively integrate features from multiple perspectives, we propose an explicitly feature enhancing strategy, Global-local unified cross-view enhancing(GLE) framework for person re-identification, based on Transformer among images sharing the same identity. Specifically, our method includes two aspects: (i) Multi-scale global feature enhancement(MSGE): It displays enhancing global features by interactively updating the features of different images sharing the same identity. (ii) Key local feature enhancement(KLFE): It aims to enhance local feature interaction in key areas by selecting and updating top-k patches based on attention scores derived from thermal maps in each training stage. Experimental results demonstrate that our method can extract more robust and discriminative features, achieving state-of-the-art performance on four pedestrian re-identification datasets. Tianran Chen, Chen Chen 0036, Xiyuan Hu |
DSAA | 2 |
| 2024 | TS-SAM: Two Small Steps for SAM, One Giant Leap for Abnormal detectionsabstractThe advent of pretrained models has significantly advanced the field of artificial intelligence. The introduction of the Segment Anything Model (SAM) has emerged as a pivotal milestone in computer vision, despite its primary application in image segmentation tasks. However, its initial performance in abnormal detection, for instance, camouflaged object detection, shadow detection and forgery detection, was not impressive. To address this, we propose to take two small steps from SAM: an efficient latent high frequency (LHF) feature enhancement module, named "LHF-Adapter", and a robust segmentation metric named "sIoU-Loss", these augmentations are seamlessly integrated into SAM, resulting in the refined model denoted as TS-SAM, tailored specifically to bolster its performance in downstream abnormal detection tasks. Extensive experiments demonstrate that the proposed method outperforms not only SAM but also the most advanced SAM variations. Our method achieves a performance improvement of up to 4.5% in camouflage object detection, setting a new state-of-the-art benchmark. Index Terms—SAM, Transformer, Adapter Dongyang Gao, Chen Chen 0036, Haotian Zhang 0025, Xiyuan Hu |
ICME | 2 |
| 2024 | Illumination Enlightened Spatial-temporal Inconsistency for Deepfake Video DetectionabstractThe rapid advancement of facial manipulation techniques has greatly simplified the creation of deepfake videos, posing a major threat to social safety, public opinions and even political stability. Existing deepfake detection methods primarily concentrate on capturing spatial artifacts or extracting uniform temporal inconsistency, neglecting the potential of exploiting dynamic spatiotemporal inconsistency. To address these issues, this paper proposes a novel network that effectively leverages dynamic spatiotemporal inconsistency, termed DSTI, by integrating the sequential illumination features and intra/inter-frame clues. The proposed DSTI contains two branches: one branch employs a transformer encoder to perform inconsistency computation from sequential illumination representations derived from 3D facial models, including illumination coefficients, 3D normal vectors, and luminance values. The other branch utilizes a timesformer network to capture intra/inter-frame inconsistency from sampled videos. Extensive experimentation validates that the proposed method outperforms other competitive approaches. Kaiyue Tian, Chen Chen 0036, Xiyuan Hu |
ICME | 2 |
| 2024 | 3D Ultrasound Image Acquisition and Diagnostic Analysis of the Common Carotid Artery with a Portable Robotic DeviceabstractUltrasound (US) imaging of the carotid artery (CA) is a non-invasive diagnostic tool widely used in the medical field to assess the condition of the carotid artery, thereby predicting the risk of cardiovascular and cerebrovascular diseases. However, implementing this method in primary healthcare can be challenging due to the requirement for professionally trained sonographers. With the adoption of US robotic devices, the probe pose can be acquired while scanning, offering the possibility for 3D reconstruction and providing analyses that are not dependent on operator experience. This article introduces a method to semi-automatically acquire serialized US images of the common carotid artery (CCA). The method involves a specially designed robotic device built with a 6-RSU parallel mechanism, which is controlled according to robot pose, force sensor data and synchronous US images. To validate the images acquired, a method is proposed to segment the intima-media of CCA and calculate the intima-media thickness (IMT), which is a key indicator for cerebrovascular events prediction. After that, we propose an algorithm to reconstruct CCA into 3D voxel data with patient movement and cardiac cycle compensated, and a longitudinal view US image of CCA can be resliced from the voxel. The methods are tested on human subjects and the results indicate that the system and workflow can provide both quantitative and qualitative information of CCA for further diagnosis. Longyue Tan, Zhaokun Deng, Mingrui Hao, Xilong Hou 0001, Chen Chen 0036, Xiaolin Gu, Xiao-Hu Zhou, Zeng-Guang Hou, Shuangyi Wang |
IROS | 6 |
| 2024 | 🔭 WATCHER: Wavelet-Guided Texture-Content Hierarchical Relation Learning for Deepfake Detection
Chen Chen 0036, Ning Zhang 0033, Xiyuan Hu |
Int. J. Comput. Vis. | 2 |
| 2024 | A new deepfake detection model for responding to perception attacks in embodied artificial intelligence
Junshuai Zheng, Xiyuan Hu, Chen Chen 0036, Dongyang Gao, Zhenmin Tang |
Image Vis. Comput. | 3 |
| 2024 | Context-aided unicity matching for person re-identification
Min Cao 0005, Cong Ding 0013, Chen Chen 0036, Silong Peng |
J. Vis. Commun. Image Represent. | 3 |
| 2024 | Distribution Unified and Probability Space Aligned Teacher-Student Learning for Imbalanced Visual RecognitionabstractImbalanced label distribution is usually the case for real-world data, which poses a challenge for training unbiased recognition model. In this paper, we study two underlying mismatches, i.e., distribution mismatch and probability space mismatch, present in class-imbalanced learning. Firstly, we analyze the label distribution mismatch between imbalanced training data and balanced test data, and introduce a distribution unified framework to unify the two distributions through probability conversion. Secondly, we analyze that the utilization of cross-entropy loss under the proposed framework may lead to probability space mismatch, where the conversion of the predictive probability is implemented in softmax probability space while the comparison with one-hot label is implemented in true probability space. To alleviate this dilemma, we involve a teacher model and formulate a teacher-student learning strategy, which contains two novel techniques. The Teacher Guided Label Smoothing (TGLS) is first proposed to relax the one-hot label to smoother pseudo softmax probability, which is more aligned with the softmax probability space. Additionally, we propose Distribution Unified Knowledge Distillation (DU-KD) under the proposed framework to further reduce both the mismatches. Experiments on several benchmarks confirm the top-level performance of the proposed method. Shaoyu Zhang 0001, Chen Chen 0036, Qiong Xie, Haigang Sun, Silong Peng |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | START: Automatic Sleep Staging with Attention-based Cross-modal Learning TransformerabstractAutomatic sleep staging is vital to scale up sleep assessment and diagnosis to serve millions experiencing sleep deprivation and disorders and enable longitudinal sleep monitoring in home environments. However, how to learn from multi-channel raw physiological signal inputs (e.g., EEG and EOG) to capture the sleep stage and physiological signal relations remains a big challenge. In this paper, we propose a sleep staging model, named Sleep Staging Cross-modal Transformer (START), which is a transformer-only method for sleep stage classification. Our model is capable of learning a joint representation from both EEG and EOG signals by using a cross-modal fusion strategy. Experimental results show that our model outperforms the state-of-the-art methods on two public datasets. Furthermore, our model provides considerable reductions in parameters and training time compared to previous methods. Jingpeng Sun, Rongxiao Wang, Gangming Zhao, Chen Chen 0036, Yixiao Qu, Xiyuan Hu, Yizhou Yu |
BIBM | 4 |
| 2023 | Dynamic Graph Learning with Content-guided Spatial-Frequency Relation Reasoning for Deepfake DetectionabstractWith the springing up of face synthesis techniques, it is prominent in need to develop powerful face forgery detection methods due to security concerns. Some existing methods attempt to employ auxiliary frequency-aware information combined with CNN backbones to discover the forged clues. Due to the inadequate information interaction with image content, the extracted frequency features are thus spatially irrelavant, struggling to generalize well on increasingly realistic counterfeit types. To address this issue, we propose a Spatial-Frequency Dynamic Graph method to exploit the relation-aware features in spatial and frequency domains via dynamic graph learning. To this end, we introduce three well-designed components: 1) Content-guided Adaptive Frequency Extraction module to mine the content-adaptive forged frequency clues. 2) Multiple Domains Attention Map Learning module to enrich the spatial-frequency contextual features with multiscale attention maps. 3) Dynamic Graph Spatial-Frequency Feature Fusion Network to explore the high-order relation of spatial and frequency features. Extensive experiments on several benchmark show that our proposed method sustainedly exceeds the state-of-the-arts by a considerable margin. Chen Chen 0036, Xiyuan Hu, Silong Peng |
CVPR | 3 |
| 2023 | Reconciling Object-Level and Global-Level Objectives for Long-Tail DetectionabstractLarge vocabulary object detectors are often faced with the long-tailed label distributions, seriously degrading their ability to detect rarely seen categories. On one hand, the rare objects are prone to be misclassified as frequent categories. On the other hand, due to the limitation on the total number of detections per image, detectors usually rank all the confidence scores globally and filter out the lower-ranking ones. This may result in missed detection during inference, especially for the rare categories that naturally come with lower scores. Existing methods mainly focus on the former problem and design various classification loss to enhance the object-level classification accuracy, but largely overlook the global-level ranking task. In this paper, we propose a novel framework that Reconciles Object-level and Global-level (ROG) objectives to address both problems. As a multi-task learning framework, ROG simultaneously trains the model with two tasks: classifying each object proposal individually and ranking all the confidence scores globally. Specifically, complementary to the object-level classification loss for model discrimination, we design a generalized average precision (GAP) loss to explicitly optimize the global-level score ranking across different objects. For each category, GAP loss generates balanced gradients to rectify the ranking errors. In experiments, we show that GAP loss is highly versatile to be plugged into various advanced methods and brings considerable benefits. Code is at https://github.com/EricZsy/ROG. Shaoyu Zhang 0001, Chen Chen 0036, Silong Peng |
ICCV | 2 |
| 2023 | RaSa: Relation and Sensitivity Aware Representation Learning for Text-based Person SearchabstractText-based person search aims to retrieve the specified person images given a textual description. The key to tackling such a challenging task is to learn powerful multi-modal representations. Towards this, we propose a Relation and Sensitivity aware representation learning method (RaSa), including two novel tasks: Relation-Aware learning (RA) and Sensitivity-Aware learning (SA). For one thing, existing methods cluster representations of all positive pairs without distinction and overlook the noise problem caused by the weak positive pairs where the text and the paired image have noise correspondences, thus leading to overfitting learning. RA offsets the overfitting risk by introducing a novel positive relation detection task (i.e., learning to distinguish strong and weak positive pairs). For another thing, learning invariant representation under data augmentation (i.e., being insensitive to some transformations) is a general practice for improving representation's robustness in existing methods. Beyond that, we encourage the representation to perceive the sensitive transformation by SA (i.e., learning to detect the replaced words), thus promoting the representation's robustness. Experiments demonstrate that RaSa outperforms existing state-of-the-art methods by 6.94%, 4.45% and 15.35% in terms of Rank@1 on CUHK-PEDES, ICFG-PEDES and RSTPReid datasets, respectively. Code is available at: https://github.com/Flame-Chasers/RaSa. Min Cao 0005, Daming Gao, Ziqiang Cao, Chen Chen 0036, Zhenfeng Fan, Liqiang Nie, Min Zhang 0005 |
IJCAI | 5 |
| 2023 | Text-based Person Search without Parallel Image-Text DataabstractText-based person search (TBPS) aims to retrieve the images of the target person from a large image gallery based on a given natural language description. Existing methods are dominated by training models with parallel image-text pairs, which are very costly to collect. In this paper, we make the first attempt to explore TBPS without parallel image-text data (μ-TBPS), in which only non-parallel images and texts, or even image-only data, can be adopted. Towards this end, we propose a two-stage framework, generation-then-retrieval (GTR), to first generate the corresponding pseudo text for each image and then perform the retrieval in a supervised manner. In the generation stage, we propose a fine-grained image captioning strategy to obtain an enriched description of the person image, which firstly utilizes a set of instruction prompts to activate the off-the-shelf pretrained vision-language model to capture and generate fine-grained person attributes, and then converts the extracted attributes into a textual description via the finetuned large language model or the hand-crafted template. In the retrieval stage, considering the noise interference of the generated texts for training model, we develop a confidence score-based training scheme by enabling more reliable texts to contribute more during the training. Experimental results on multiple TBPS benchmarks (i.e., CUHK-PEDES, ICFG-PEDES and RSTPReid) show that the proposed GTR can achieve a promising performance without relying on parallel image-text data. Min Cao 0005, Chen Chen 0036, Ziqiang Cao, Liqiang Nie, Min Zhang 0005 |
ACM Multimedia | 4 |
| 2023 | Mining Temporal Inconsistency with 3D Face Model for Deepfake Video Detection
Ziyi Cheng, Chen Chen 0036, Xiyuan Hu |
PRCV (7) | 2 |
| 2023 | Balanced knowledge distillation for long-tailed learning
Shaoyu Zhang 0001, Chen Chen 0036, Xiyuan Hu, Silong Peng |
Neurocomputing | 2 |
| 2023 | A landmark-free approach for automatic, dense and robust correspondence of 3D faces
Zhenfeng Fan, Xiyuan Hu, Chen Chen 0036, Xiaolian Wang, Silong Peng |
Pattern Recognit. | 3 |
| 2023 | Causality Network of Infectious Disease Revealed With Causal DecompositionabstractCausal inference in the field of infectious disease attempts to gain insight into the potential causal nature of an association between risk factors and diseases. Simulated causality inference experiments have shown preliminary promise in improving understanding of the transmission of infectious diseases but still lack sufficient quantitative causal inference studies based on real-world data. Here, we investigate the causal interactions between three different infectious diseases and related factors, using causal decomposition analysis, to characterize the nature of infectious disease transmission. We show that the complex interactions between infectious disease and human behavior have a quantifiable impact on transmission efficiency of infectious diseases. Our findings, by shedding light on the underlying transmission mechanism of infectious diseases, suggest that causal inference analysis is a promising approach to determine epidemiological interventions. Jingpeng Sun, Chen Chen 0036, Hesong Wang, Yuxing Zhi, Silong Peng, Chung-Kang Peng, Norden E. Huang, Guangrui Huang, Albert Yang |
IEEE J. Biomed. Health Informatics | 3 |
| 2023 | Progressive Context-Aware Graph Feature Learning for Target Re-IdentificationabstractThis paper aims at robust and discriminative feature learning for target re-identification (Re-ID). In addition to paying attention to the individual appearance information as in most Re-ID methods, we further utilize the abundant contextual information as additional clues to guide the feature learning. Graph as a format of structured data is used to represent the target sample with its context. It describes the first-order appearance information of the samples and the second-order topological relationship information among samples, based on which we compute the feature representation by learning a graph feature embedding. We provide a detailed analysis of graph convolutional network mechanism applied in target Re-ID and propose a novel progressive context-aware graph feature learning method, in which the message passing is dominated by a pre-defined adjacency relationship followed by a learned relationship in a self-adaptive way. The proposed method fully exploits and utilizes contextual information at a low cost for Re-ID. Extensive experiments on five Re-ID benchmarks demonstrate the state-of-the-art performance of the proposed method. Min Cao 0005, Cong Ding 0013, Chen Chen 0036, Hao Dou, Xiyuan Hu, Junchi Yan |
IEEE Trans. Multim. | 3 |
| 2022 | Improving Surveillance Object Detection with Adaptive Omni-Attention over Both Inter-frame and Intra-frame Context
Chen Chen 0036, Xiyuan Hu |
ACCV (2) | 2 |
| 2022 | Unstructured Feature Decoupling for Vehicle Re-identification
Hao Luo 0004, Silong Peng, Fan Wang 0019, Chen Chen 0036, Hao Li 0030 |
ECCV (14) | 5 |
| 2022 | Label-Occurrence-Balanced Mixup for Long-Tailed RecognitionabstractMixup is a popular data augmentation method, with many variants subsequently proposed. These methods mainly create new examples via convex combination of random data pairs and their corresponding one-hot labels. However, most of them adhere to a random sampling and mixing strategy, without considering the frequency of label occurrence in the mixing process. When applying mixup to long-tailed data, a label suppression issue arises, where the frequency of label occurrence for each class is imbalanced and most of the new examples will be completely or partially assigned with head labels. The suppression effect may further aggravate the problem of data imbalance and lead to a poor performance on tail classes. To address this problem, we propose Label-Occurrence-Balanced Mixup to augment data while keeping the label occurrence for each class statistically balanced. In a word, we employ two independent class-balanced samplers to select data pairs and mix them to generate new data. We test our method on several long-tailed vision and sound recognition benchmarks. Experimental results show that our method significantly promotes the adaptability of mixup method to imbalanced data and achieves superior performance compared with state-of-the-art long-tailed learning methods. Shaoyu Zhang 0001, Chen Chen 0036, Silong Peng |
ICASSP | 2 |
| 2022 | Partner learning: A comprehensive knowledge transfer for vehicle re-identification
Zhiqun He, Chen Chen 0036, Silong Peng |
Neurocomputing | 3 |
| 2022 | Corrigendum to "Partner learning: A comprehensive knowledge transfer for vehicle re-identification" [Neurocomputing 480 (2022) 89-98/NEUCOM-D-21-03435R1]
Zhiqun He, Chen Chen 0036, Silong Peng |
Neurocomputing | 3 |
| 2022 | Regularizing deep networks with label geometry for accurate object localization on small training datasets
Xiaolian Wang, Xiyuan Hu, Chen Chen 0036, Silong Peng |
Pattern Recognit. Lett. | 3 |
| 2022 | Navigating Diverse Salient Features for Vehicle Re-IdentificationabstractMining sufficient discriminative information is vital for effective feature representation in vehicle re-identification. Traditional methods mainly focus on the most salient features and neglect whether the explored information is sufficient. This paper tackles the above limitation by proposing a novel Salience-Navigated Vehicle Re-identification Network (SVRN) which explores diverse salient features at multi-scales. For mining sufficient salient features, we design SVRN from two aspects: 1) network architecture: we propose a novel salience-navigated vehicle re-identification network, which mines diverse features under a cascaded suppress-and-explore mode. 2) feature space: cross-space constraint enables the diversity from feature space, which restrains the cross-space features by vehicle and image identifications (IDs). Extensive experiments demonstrate our method’s effectiveness, and the overall results surpass all previous state-of-the-arts in three widely-used Vehicle ReID benchmarks (VeRi-776, VehicleID, and VERI-WILD), i.e., we achieve an 84.5% mAP on VeRi-776 benchmark that outperforms the second-best method by a large margin (3.5% mAP). Zhiqun He, Chen Chen 0036, Silong Peng |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Pseudo Graph Convolutional Network for Vehicle ReIDabstractImage-based Vehicle ReID methods have suffered from limited information caused by viewpoints, illumination, and occlusion as they usually use a single image as input. Graph convolutional methods (GCN) can alleviate the aforementioned problem by aggregating neighbor samples' information to enhance the feature representation. However, it's uneconomical and computational for the inference processes of GCN-based methods since they need to iterate over all samples for searching the neighbor nodes. In this paper, we propose the first Pseudo-GCN Vehicle ReID method (PGVR) which enables a CNN-based module to performs competitively to GCN-based methods and has a faster and lightweight inference process. To enable the Pseudo-GCN mechanism, a two-branch network and a graph-based knowledge distillation are proposed. The two-branch network consists of a CNN-based student branch and a GCN-based teacher branch. The GCN-based teacher branch adopts a ReID-based GCN to learn the topological optimization ability under the supervision of ReID tasks during training time. Moreover, the graph-based knowledge distillation explicitly transfers the topological optimization ability from the teacher branch to the student branch which acknowledges all nodes. We evaluate our proposed method PGVR on three mainstream Vehicle ReID benchmarks and demonstrate that PGVR achieves state-of-the-art performance. Zhiqun He, Silong Peng, Chen Chen 0036, Wei Wu 0021 |
ACM Multimedia | 4 |
| 2021 | Bioinformatics and machine learning methodologies to identify the effects of central nervous system disorders on glioblastoma progressionabstractGlioblastoma (GBM) is a common malignant brain tumor which often presents as a comorbidity with central nervous system (CNS) disorders. Both CNS disorders and GBM cells release glutamate and show an abnormality, but differ in cellular behavior. So, their etiology is not well understood, nor is it clear how CNS disorders influence GBM behavior or growth. This led us to employ a quantitative analytical framework to unravel shared differentially expressed genes (DEGs) and cell signaling pathways that could link CNS disorders and GBM using datasets acquired from the Gene Expression Omnibus database (GEO) and The Cancer Genome Atlas (TCGA) datasets where normal tissue and disease-affected tissue were examined. After identifying DEGs, we identified disease-gene association networks and signaling pathways and performed gene ontology (GO) analyses as well as hub protein identifications to predict the roles of these DEGs. We expanded our study to determine the significant genes that may play a role in GBM progression and the survival of the GBM patients by exploiting clinical and genetic factors using the Cox Proportional Hazard Model and the Kaplan-Meier estimator. In this study, 177 DEGs with 129 upregulated and 48 downregulated genes were identified. Our findings indicate new ways that CNS disorders may influence the incidence of GBM progression, growth or establishment and may also function as biomarkers for GBM prognosis and potential targets for therapies. Our comparison with gold standard databases also provides further proof to support the connection of our identified biomarkers in the pathology underlying the GBM progression. Humayan Kabir Rana, Silong Peng, Xiyuan Hu, Chen Chen 0036, Julian M. W. Quinn, Mohammad Ali Moni |
Briefings Bioinform. | 5 |
| 2021 | Progressive Bilateral-Context Driven Model for Post-Processing Person Re-IdentificationabstractMost existing person re-identification methods compute pairwise similarity by extracting robust visual features and learning the discriminative metric. Owing to visual ambiguities, these content-based methods that determine the pairwise relationship only based on the similarity between them, inevitably produce a suboptimal ranking list. Instead, the pairwise similarity can be estimated more accurately along the geodesic path of the underlying data manifold by exploring the rich contextual information of the sample. In this paper, we propose a lightweight post-processing person re-identification method in which the pairwise measure is determined by the relationship between the sample and the counterpart's context in an unsupervised way. We translate the point-to-point comparison into the bilateral point-to-set comparison. The sample's context is composed of its neighbor samples with two different definition ways: the first order context and the second order context, which are used to compute the pairwise similarity in sequence, resulting in a progressive post-processing model. The experiments on four large-scale person re-identification benchmark datasets indicate that (1) the proposed method can consistently achieve higher accuracies by serving as a post-processing procedure after the content-based person re-identification methods, showing its state-of-the-art results, (2) the proposed lightweight method only needs about 6 milliseconds for optimizing the ranking results of one sample, showing its high-efficiency. Code is available at: https://github.com/123ci/PBCmodel. Min Cao 0005, Chen Chen 0036, Hao Dou, Xiyuan Hu, Silong Peng, Arjan Kuijper |
IEEE Trans. Multim. | 2 |
| 2020 | Illuminating Vehicles With Motion Priors For Surveillance Vehicle DetectionabstractVehicle detection in traffic surveillance videos is a special subtask in object detection, where desired objects are vehicles moving on the road while the background is still within a sequence. The disparity of speed within each frame, i.e. moving and static, is consistent with the vehicle and background semantic to some extent, thus motions can be extracted to enhance the appearance of foreground. In this paper, we propose a motion prior embedded parallel architecture for vehicle detection, aiming at illuminating vehicles and suppressing false positives in the background. We further implement extensive experiments on the UA-DETRAC dataset to validate the effectiveness of our approach, and achieve promising performance in both accuracy and speed. Xiaolian Wang, Xiyuan Hu, Chen Chen 0036, Zhenfeng Fan, Silong Peng |
ICIP | 3 |
| 2020 | Deep Top-rank Counter Metric for Person Re-identificationabstractIn the research field of person re-identification, deep metric learning that guides the efficient and effective embedding learning serves as one of the most fundamental tasks. Recent efforts of the loss function based deep metric learning methods mainly focus on the top rank accuracy optimization by minimizing the distance difference between the correctly matching sample pair and wrongly matched sample pair. However, it is more straightforward to count the occurrences of correct top-rank candidates and maximize the counting results for better top rank accuracy. In this paper, we propose a generalized logistic function based metric with effective practicalness in deep learning, namely the“deeptop-rankcountermetric”, to approximately optimize the counted occurrences of the correct top-rank matches. The properties that qualify the proposed metric as a well-suited deep re-identification metric have been discussed and a progressive hard sample mining strategy is also introduced for effective training and performance boosting. The extensive experiments show that the proposed top-rank counter metric outperforms other loss function based deep metrics and achieves the state-of-the-art accuracies. Chen Chen 0036, Hao Dou, Xiyuan Hu, Silong Peng |
ICPR | 1 |
| 2020 | Towards Low-Bit Quantization of Deep Neural Networks with Limited DataabstractRecent machine learning methods use increasingly large deep neural networks to achieve state-of-the-art results in various tasks. Network quantization can effectively reduce computation and memory costs without modifying network structures, facilitating the deployment of deep neural networks (DNNs) on cloud and edge devices. However, most of the existing methods usually need time-consuming training or fine-tuning and access to the original training dataset that may be unavailable due to privacy or security concerns. In this paper, we present a novel method to achieve low-precision quantization with limited data. Firstly, to reduce the complexity of per-channel quantization and degeneration of per-layer quantization, we introduce group quantization that separates the output channels into groups and processes each group independently. Secondly, to better distill knowledge from the pre-trained FP32 model with limited data, we introduce a two-stage knowledge distillation method that divides the optimization process into blockwise optimization and joint optimization to address the limitation of layer-wise supervision and global supervision. Extensive experiments on ImageNet2012 (ResNet18/50, ShuffleNetV2, and MobileNetV2) demonstrate that the proposed approach can significantly improve the quantization model's accuracy when only a few training samples are available. We further show that the method also extends to other computer vision architectures and tasks such as object detection. Chen Chen 0036, Xiyuan Hu, Silong Peng |
ICPR | 2 |
| 2020 | EvoQ: Mixed Precision Quantization of DNNs via Sensitivity Guided Evolutionary SearchabstractNetwork quantization can effectively reduce computation and memory costs without modifying network structures, facilitating the deployment of deep neural networks (DNNs) on edge devices. However, most of the existing methods usually need time-consuming training or fine-tuning and access to the original training dataset that may be unavailable due to privacy or security concerns. In this paper, we introduce a novel method named EvoQ that employs evolutionary search to achieve mixed precision quantization with limited data, which can optimize the resource allocation without adding computation consumption. Considering the shortage of samples and expensive search costs, we use 50 samples to measure the output difference between the quantization model and the pre-trained model for the evaluation of quantization policy, which can save the time obviously while maintaining high accuracy. To improve the search efficiency, we analyze the quantization sensitivity of each layer and utilize the results to optimize the mutation operation. At last, we calibrate the outputs and intermediate features of the quantization model using the selected 50 samples to improve the performance further. We implement extensive experiments on a diverse set of models, including ResNet18/50/101, SqueezeNet, ShuffleNetV2, and MobileNetV2 on ImageNet, as well as SSD-VGG and SSD-ResNet50 on PASCAL VOC. Our method can improve the performance apparently and outperforms the existing post-training quantization methods, demonstrating the effectiveness of EvoQ. Chen Chen 0036, Xiyuan Hu, Silong Peng |
IJCNN | 2 |
| 2020 | PCA-SRGAN: Incremental Orthogonal Projection Discrimination for Face Super-resolutionabstractGenerative Adversarial Networks (GANs) have been employed for face super resolution but they bring distorted facial details easily and still have weakness on recovering realistic texture. To further improve the performance of GAN-based models on super-resolving face images, we propose PCA-SRGAN which pays attention to the cumulative discrimination in the orthogonal projection space spanned by PCA projection matrix of face data. By feeding the principal component projections ranging from structure to details into the discriminator, the discrimination difficulty will be greatly alleviated and the generator can be enhanced to reconstruct clearer contour and finer texture, helpful to achieve the high perception and low distortion eventually. This incremental orthogonal projection discrimination has ensured a precise optimization procedure from coarse to fine and avoids the dependence on the perceptual regularization. We conduct experiments on CelebA and FFHQ face datasets. The qualitative visual effect and quantitative evaluation have demonstrated the overwhelming performance of our model over related works. Hao Dou, Chen Chen 0036, Xiyuan Hu, Zuxing Xuan, Zhisen Hu, Silong Peng |
ACM Multimedia | 2 |
| 2020 | Asymmetric CycleGAN for image-to-image translations with uneven complexities
Hao Dou, Chen Chen 0036, Xiyuan Hu, Libang Jia, Silong Peng |
Neurocomputing | 2 |
| 2019 | Boosting Local Shape Matching for Dense 3D Face CorrespondenceabstractDense 3D face correspondence is a fundamental and challenging issue in the literature of 3D face analysis. Correspondence between two 3D faces can be viewed as a non-rigid registration problem that one deforms into the other, which is commonly guided by a few facial landmarks in many existing works. However, the current works seldom consider the problem of incoherent deformation caused by landmarks. In this paper, we explicitly formulate the deformation as locally rigid motions guided by some seed points, and the formulated deformation satisfies coherent local motions everywhere on a face. The seed points are initialized by a few landmarks, and are then augmented to boost shape matching between the template and the target face step by step, to finally achieve dense correspondence. In each step, we employ a hierarchical scheme for local shape registration, together with a Gaussian reweighting strategy for accurate matching of local features around the seed points. In our experiments, we evaluate the proposed method extensively on several datasets, including two publicly available ones: FRGC v2.0 and BU-3DFE. The experimental results demonstrate that our method can achieve accurate feature correspondence, coherent local shape motion, and compact data representation. These merits actually settle some important issues for practical applications, such as expressions, noise, and partial data. Zhenfeng Fan, Xiyuan Hu, Chen Chen 0036, Silong Peng |
CVPR | 3 |
| 2019 | Asymmetric Cyclegan for Unpaired NIR-to-RGB Face Image TranslationabstractTranslating near-infrared (NIR) face into color (RGB) face, is helpful to improve the visual effect of images and the performance of face recognition. The model for unpaired image-to-image translation is suitable for this task due to the high cost of pixel-matched data. Because of the complexity difference between NIR and RGB image domains, the complexity inequality in bidirectional NIR-RGB translations is significant. We analyze the limitation of the original CycleGAN in asymmetric translation tasks, and propose an Asymmetric Cycle-GAN model with U-net-like generators of unequal sizes to adapt to the asymmetric need in NIR-RGB translations. The edge-retain loss between NIR and the generated RGB images is also introduced to enhance face visual quality. The qualitative visual evaluation and quantitative evaluation with face ID and skin color criteria show that our model achieves great improvements compared with state-of-the-art methods on three public datasets and a newly proposed dataset. Hao Dou, Chen Chen 0036, Xiyuan Hu, Silong Peng |
ICASSP | 2 |
| 2019 | Improving Object Detection with Consistent Negative Sample Mining
Xiaolian Wang, Xiyuan Hu, Chen Chen 0036, Zhenfeng Fan, Silong Peng |
ICONIP (2) | 3 |
| 2019 | TP-ADMM: An Efficient Two-Stage Framework for Training Binary Neural Networks
Chen Chen 0036, Xiyuan Hu, Silong Peng |
ICONIP (4) | 2 |
| 2019 | Towards fast and kernelized orthogonal discriminant analysis on person re-identification
Min Cao 0005, Chen Chen 0036, Xiyuan Hu, Silong Peng |
Pattern Recognit. | 2 |
| 2018 | Ranking Loss: A Novel Metric Learning Method for Person Re-identification
Min Cao 0005, Chen Chen 0036, Xiyuan Hu, Silong Peng |
ACCV (2) | 2 |
| 2018 | Dense Semantic and Topological Correspondence of 3D Faces without Landmarks
Zhenfeng Fan, Xiyuan Hu, Chen Chen 0036, Silong Peng |
ECCV (16) | 3 |
| 2018 | Region-specific Metric Learning for Person Re-identificationabstractPerson re-identification addresses the problem of matching individual images of the same person captured by different non-overlapping camera views. Distance metric learning plays an effective role in addressing the problem. With the features extracted on several regions of person image, most of distance metric learning methods have been developed in which the learnt cross-view transformations are region-generic, i.e all region-features share a homogeneous transformation. The spatial structure of person image is ignored and the distribution difference among different region-features is neglected. Therefore in this paper, we propose a novel region-specific metric learning method in which a series of region-specific sub-models are optimized for learning cross-view region-specific transformations. Additionally, we also present a novel feature pre-processing scheme that is designed to improve the features' discriminative power by removing weakly discriminative features. Experimental results on the publicly available VIPeR, PRID450S and QMUL GRID datasets demonstrate that the proposed method performs favorably against the state-of-the-art methods. Min Cao 0005, Chen Chen 0036, Xiyuan Hu, Silong Peng |
ICPR | 2 |
| 2017 | Key Person Aided Re-identification in Partially Ordered Pedestrian Set
Chen Chen 0036, Min Cao 0005, Silong Peng |
BMVC | 1 |
| 2016 | Face spoofing detection based on 3D lighting environment analysis of image pairabstractIn this paper, we present a novel face spoofing detection method based on 3D lighting environment analysis of an image pair collected before and after the lighting environment change. Our idea is inspired from the unimpressive fact that the illumination distributions of the internal spoof face stays stable under the protection of the photo and screen plane, while that of a exposed genuine face changes accordingly to different lighting environment due to a natural response of 3D structure. After estimating two sets of lighting environment coefficients of client's face image pair with the hand of 3D Morphable Model (3DMM) and Sphere Harmonic Illumination Model (SHIM), robust liveness judgement is conducted by hypothesis tests. Experimental results show the effectiveness of proposed method on multiple kinds of face attacks including printed photo, screen photo, and video replay attack, and other advantages such as user cooperation free, loose using conditions, simple equipment demand, easy to camouflage and propitious to face recognition. Xiyuan Hu, Chen Chen 0036, Silong Peng |
ICPR | 4 |
| 2016 | Emotion in Context: Deep Semantic Feature Fusion for Video Emotion RecognitionabstractHuman emotions demonstrate high correlations with certain events that consist of the interaction of objects, and are usually constrained by particular scenes. In this paper, we exploit these abundant context clues for video emotion recognition. We first compute event, object and scene scores with state-of-the-art detectors based on deep neural networks. The extracted high-level features serve as effective contextual information demonstrating what is occurring in the video, and are further integrated in a context fusion network to generate a unified representation aiming to bridge the affective gap. The contribution of this paper is three-fold: (a) we are the first to incorporate event as context for emotion recognition, and we demonstrate its superiority for emotion understanding, (b) we utilize a context fusion network to exploit a comprehensive set of high-level semantics features as contextual clues and (c) the proposed framework can realize recognition in real-time and achieves state-of-the-art performance on two challenging emotion recognition benchmarks, 50.6 on VideoEmotion and 51.8 on Ekman. Chen Chen 0036, Zuxuan Wu, Yu-Gang Jiang 0001 |
ACM Multimedia | 1 |