VLDB 2026 Research / reviewers in the wild / expert
Dapeng Chen
dblp:04/3068
· DBLP profile ↗
70ranked-venue papers
14as first author
33since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 52 · 11 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 44 · 10 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 3Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The Evolution of Philosophy: A Metaphorical Cognition Perspective
Rui Mao 0010, Dapeng Chen, Xulang Zhang, Erik Cambria |
LREC | 2 |
| 2026 | Uncertainty-constrained fusion of single-view and multi-view depth estimation for AR virtual-real occlusion
Shuai Ding 0001, Yongze Li, Lina Wei, Dapeng Chen |
Neural Networks | 6 |
| 2025 | Few-Shot Incremental Multi-modal Learning via Touch Guidance and Imaginary Vision SynthesisabstractMultimodal perception, which integrates vision and touch, is increasingly demonstrating its significance in domains such as embodied intelligence and human-computer interaction. However, in open-world scenarios, multimodal data streams face significant challenges, including catastrophic forgetting and overfitting, during few-shot class incremental learning (FSCIL), leading to a severe degradation in model performance. In this work, we propose a novel approach named Few-Shot Incremental Multi-modal Learning via Touch Guidance and Imaginary Vision Synthesis (TIFS). Our method leverages vision imagination synthesis to enhance the semantic understanding and integrates touch and vision fusion to improve the problem of modal imbalance. Specifically, we introduce a framework that employs touch-guided vision information for cross-modal contrastive learning to address the challenges of few-shot learning. Additionally, we incorporate multiple learning mechanisms, including regularization, memory mechanisms, and attention mechanisms, to mitigate catastrophic forgetting during multi-incremental step learning. Experimental results on the Touch and Go and VisGel datasets demonstrate that the TIFS framework exhibits robust continuous learning capabilities and strong generalization performance in touch-vision few-shot incremental learning tasks. Our code is available at https://github.com/Vision-Multimodal-Lab-HZCU/TIFS. Lina Wei, Zhongsheng Lin, Canghong Jin, Hanbin Zhao, Dapeng Chen |
IJCAI | 7 |
| 2025 | Joint Geometric Self-Attention and Boundary-Aware Search for High-Precision Intracranial Aneurysm Mesh Segmentation
Fuhao Zhang, Ling Wang 0005, Dapeng Chen, Jinshan Tang, Jingfeng Jiang, Nan Mu |
SMC | 6 |
| 2025 | A Transformer-Based Dual-Branch Mesh Convolutional Neural Network for Aortic Dissection SegmentationabstractAortic dissection (AD) is a life-threatening condition caused by a tear in the aortic intima, allowing blood to enter the vessel wall and form a false lumen. Due to its high mortality rate, timely diagnosis and precise treatment are critical. Clinical diagnosis and treatment of AD rely heavily on accurate 3D vascular image segmentation. To address existing methods’ low segmentation accuracy and insufficient geometric detail preservation, this paper proposes a Transformer-based Dual-Branch Mesh Segmentation Network (TD-MSeg) for AD. This network employs a mesh-based self-attention mechanism to retain vascular geometric details while adopting a dual-branch decoder to effectively fuse features and model long-range dependencies. Specifically, TD-MSeg incorporates three key components: a Hierarchical Mesh Transformer (HMT) module that enhances feature modeling of critical anatomical structures (e.g., intimal tears), a dual-branch decoder that facilitates collaborative optimization of multi-scale local and global features, and a mesh label refinement module that uses a wide-path exploration algorithm to eliminate deformation artifacts and improve spatial label continuity. Moreover, experiments on two AD mesh segmentation datasets demonstrate that the proposed TD-MSeg achieves a 6% improvement in accuracy compared to traditional models and significantly enhances the recognition of complex vascular structures, thereby providing high-precision 3D reconstruction support for endovascular surgical planning. Fuhao Zhang, Ling Wang 0005, Dapeng Chen, Jinshan Tang, Jingfeng Jiang, Nan Mu |
SMC | 6 |
| 2025 | High-resolution multi-view stereo with multi-scale feature fusionabstractTo enhance the handling of three-dimensional reconstruction for large-scale scenes and high-resolution images, we introduce a novel multi-view high-resolution three-dimensional reconstruction approach. Our proposed method integrates a Feature Pyramid Network with the Swin Transformer for improved performance. We integrate the Swin Transformer into the feature pyramid. This integration aims to establish long-range feature dependencies, facilitate information exchange between different input positions, and enhance the global consistency of feature representation. This improves the efficiency of the feature extraction stage. Following this, we apply cost volume regularization to mitigate noise and compute depth maps. A Depth Optimization Module is employed to refine the predicted depth maps, thereby enhancing their precision. Experimental results demonstrate the efficacy of our method in generating more accurate depth information, particularly in predicting high-resolution depth maps. Our approach utilizes these depth predictions to generate point clouds, enabling precise matching and reconstruction of multi-view images. Experiments conducted on public datasets validate the effectiveness and superiority of our proposed method. Dapeng Chen, Nanxuan Huang, Jia Liu 0034 |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | Low-delay haptic texture display method based on user action information and texture image
Dapeng Chen, Tianyu Fan, Jia Liu 0034, Aiguo Song |
Int. J. Hum. Comput. Stud. | 1 |
| 2025 | Adaptive Human Movement Compensation Control of Supernumerary Robotic Limb for Overhead Support Task With Non-Zero-Sum Differential Game TheoryabstractThe supernumerary robotic limb (SRL) mounted on the shoulder has been demonstrated to be able to serve as a third arm to assist human in overhead support task. However, the mechanical connection between the wearer and the SRL means that human operator’s movement will continually disturb the SRL and may lead to instability. Moreover, there may be physical conflicts between SRL and human operator due to their different intentions, which may potentially increase the load on human operator. Therefore, it’s necessary to control SRL to ensure stable supporting and transparent interaction. Here, we propose an adaptive controller for human movement compensation. Firstly, we model the human-SRL coordinative behavior based on non-zero-sum differential game theory aiming to enhance the support stability and reduce the operator’s load, which is a framework capable of dynamically regulating the control strategies between two interacting agents. We then implement an adaptive control strategy that adjusts the SRL’s input optimally responsive to the human operator’s input in the sense of Nash equilibrium to meet predefined control objectives. From experimental results during overhead support task, the proposed controller reduces the peak reaction force from 15.59 N to 4.50 N compared to the state-of-the-art controller that disregards human input while ensuring the stability of support, thereby providing advantage for human-SRL coordination in overhead support task. Jianxi Zhang, Hong Zeng 0001, Jia Liu 0034, Dapeng Chen, Aiguo Song |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | Design and Evaluation of a 6-DoF Wearable Fingertip Device for Haptic Shape RenderingabstractAs virtual objects contain increasingly rich attribute information, small wearable fingertip devices need to have higher degrees of freedom (DoFs) to convey the haptic sensation of virtual objects. In order to effectively display the shape features of virtual objects to users through curvature, we designed a 6-DoF wearable fingertip device (WFD). This WFD combines a 6-DoF Stewart parallel mechanism, consisting of a static platform and a mobile platform connected by six revolute-spherical-spherical kinematic chains. The translation and rotation of the mobile platform are driven by six miniature servo motors, which can simulate haptic sensations such as making and breaking contact, sliding, and skin stretch when the fingertip interacts with a virtual surface. The WFD is fixed at the user's dominant index finger using hook-and-loop fasteners, with a size of 68 × 59 × 56 mm$^{3}$3 and a mass of 45.5 g. We analyzed and validated the kinematic model of the WFD and tested its force output capability. Finally, we invited 15 adults to conduct three subjective perception experiments to evaluate the performance of the WFD in curvature perception and shape display. The experimental results show that: (1) The just noticeable difference (JND) for curvature identification using the WFD is 3.02$\pm$±0.23 m$^{-1}$-1; (2) The 6-DoF haptic feedback provided by the WFD improves the accuracy of curved surface recognition from 53.4$\pm$±7.1% in 3-DoF to 72.0$\pm$±5.9%; (3) Even without visual feedback, the shape recognition accuracy of the WFD when combined with the Touch device reaches 82.3$\pm$±8.2% . Experimental results show that the WFD has good performance and potential in curvature perception and shape display. Dapeng Chen, Haojun Ni, Lifeng Zhu, Hong Zeng 0001, Jia Liu 0034, Aiguo Song |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2024 | Align Before Adapt: Leveraging Entity-to-Region Alignments for Generalizable Video Action RecognitionabstractLarge-scale visual-language pre-trained models have achieved significant success in various video tasks. However, most existing methods follow an “adapt then align” paradigm, which adapts pre-trained image encoders to model video-level representations and utilizes one-hot or text embedding of the action labels for supervision. This paradigm overlooks the challenge of mapping from static images to complicated activity concepts. In this paper, we propose a novel “Align before Adapt” (ALT) paradigm. Prior to adapting to video representation learning, we exploit the entity-to-region alignments for each frame. The alignments are fulfilled by matching the region-aware image embeddings to an offline-constructed text corpus. With the aligned entities, we feed their text embeddings to a transformer-based video adapter as the queries, which can help extract the semantics of the most important entities from a video to a vector. This paradigm reuses the visual-language alignment of VLP during adaptation and tries to explain an action by the underlying entities. This helps understand actions by bridging the gap with complex activity semantics, particularly when facing unfamiliar or unseen categories. ALT demonstrates competitive performance while maintaining remarkably low computational costs. In fully supervised experiments, it achieves 88.1 % top-1 accuracy on Kinetics-400 with only 4947 GFLOPs. Moreover, ALT outperforms the previous state-of-the-art methods in both zero-shot and fewshot experiments, emphasizing its superior generalizability across various learning scenarios. Yifei Chen 0010, Dapeng Chen, Ruijin Liu, Wenyuan Xue, Wei Peng 0011 |
CVPR | 2 |
| 2024 | A Neuroinspired Contrast Mechanism enables Few-Shot Object Detection
Lingxiao Yang, Dapeng Chen, Yifei Chen 0010, Wei Peng 0011, Xiaohua Xie |
Pattern Recognit. | 2 |
| 2024 | Structured Domain Adaptation With Online Relation Regularization for Unsupervised Person Re-IDabstractUnsupervised domain adaptation (UDA) aims at adapting the model trained on a labeled source-domain dataset to an unlabeled target-domain dataset. The task of UDA on open-set person reidentification (re-ID) is even more challenging as the identities (classes) do not have overlap between the two domains. One major research direction was based on domain translation, which, however, has fallen out of favor in recent years due to inferior performance compared with pseudo-label-based methods. We argue that domain translation has great potential on exploiting valuable source-domain data but the existing methods did not provide proper regularization on the translation process. Specifically, previous methods only focus on maintaining the identities of the translated images while ignoring the intersample relations during translation. To tackle the challenges, we propose an end-to-end structured domain adaptation framework with an online relation-consistency regularization term. During training, the person feature encoder is optimized to model intersample relations on-the-fly for supervising relation-consistency domain translation, which in turn improves the encoder with informative translated images. The encoder can be further improved with pseudo labels, where the source-to-target translated images with ground-truth identities and target-domain images with pseudo identities are jointly used for training. In the experiments, our proposed framework is shown to achieve state-of-the-art performance on multiple UDA tasks of person re-ID. With the synthetic→real translated images from our structured domain-translation network, we achieved second place in the Visual Domain Adaptation Challenge (VisDA) in 2020. Yixiao Ge, Feng Zhu 0006, Dapeng Chen, Rui Zhao 0001, Xiaogang Wang 0001, Hongsheng Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Improving Table Structure Recognition with Visual-Alignment Sequential Coordinate ModelingabstractTable structure recognition aims to extract the logical and physical structure of unstructured table images into a machine-readable format. The latest end-to-end image-to-text approaches simultaneously predict the two structures by two decoders, where the prediction of the physical structure (the bounding boxes of the cells) is based on the representation of the logical structure. However, the previous methods struggle with imprecise bounding boxes as the logical representation lacks local visual information. To address this issue, we propose an end-to-end sequential modeling framework for table structure recognition called VAST. It contains a novel coordinate sequence decoder triggered by the representation of the non-empty cell from the logical structure decoder. In the coordinate sequence decoder, we model the bounding box coordinates as a language sequence, where the left, top, right and bottom coordinates are decoded sequentially to leverage the inter-coordinate dependency. Furthermore, we propose an auxiliary visual-alignment loss to enforce the logical representation of the non-empty cells to contain more local visual details, which helps produce better cell bounding boxes. Extensive experiments demonstrate that our proposed method can achieve state-of-the-art results in both logical and physical structure recognition. The ablation study also validates that the proposed coordinate sequence decoder and the visual-alignment loss are the keys to the success of our method. Yongshuai Huang, Ning Lu 0003, Dapeng Chen, Zecheng Xie, Shenggao Zhu, Liangcai Gao, Wei Peng 0011 |
CVPR | 3 |
| 2023 | Video Action Recognition with Attentive Semantic UnitsabstractVisual-Language Models (VLMs) have significantly advanced video action recognition. Supervised by the semantics of action labels, recent works adapt the visual branch of VLMs to learn video representations. Despite the effectiveness proved by these works, we believe that the potential of VLMs has yet to be fully harnessed. In light of this, we exploit the semantic units (SU) hiding behind the action labels and leverage their correlations with fine-grained items in frames for more accurate action recognition. SUs are entities extracted from the language descriptions of the entire action set, including body parts, objects, scenes, and motions. To further enhance the alignments between visual contents and the SUs, we introduce a multi-region attention module (MRA) to the visual branch of the VLM. The MRA allows the perception of region-aware visual features beyond the original global feature. Our method adaptively attends to and selects relevant SUs with visual features of frames. With a cross-modal decoder, the selected SUs serve to decode spatiotemporal video representations. In summary, the SUs as the medium can boost discriminative ability and transferability. Specifically, in fully-supervised learning, our method achieved 87.8% top-1 accuracy on Kinetics-400. In K=2 few-shot experiments, our method surpassed the previous state-of-the-art by +7.1% and +15.0% on HMDB-51 and UCF-101, respectively. Yifei Chen 0010, Dapeng Chen, Ruijin Liu, Wei Peng 0011 |
ICCV | 2 |
| 2023 | Video Summarization Leveraging Multimodal Information for Presentations
Dapeng Chen, Rongjun Li, Wenyuan Xue, Wei Peng 0011 |
INTERSPEECH | 2 |
| 2023 | PBFormer: Capturing Complex Scene Text Shape with Polynomial Band TransformerabstractWe present PBFormer, an efficient yet powerful scene text detector that unifies the transformer with a novel text shape representation Polynomial Band (PB). The representation has four polynomial curves to fit a text's top, bottom, left, and right sides, which can capture a text with a complex shape by varying polynomial coefficients. PB has appealing features compared with conventional representations: 1) It can model different curvatures with a fixed number of parameters, while polygon-points-based methods need to utilize a different number of points. 2) It can distinguish adjacent or overlapping texts as they have apparent different curve coefficients, while segmentation-based or points-based methods suffer from adhesive spatial positions. PBFormer combines the PB with the transformer, which can directly generate smooth text contours sampled from predicted curves without interpolation. A parameter-free cross-scale pixel attention (CPA) module is employed to highlight the feature map of a suitable scale while suppressing the other feature maps. The simple operation can help detect small-scale texts and is compatible with the one-stage DETR framework, where no postprocessing exists for NMS. Furthermore, PBFormer is trained with a shape-contained loss, which not only enforces the piecewise alignment between the ground truth and the predicted curves but also makes curves' position and shapes consistent with each other. Without bells and whistles about text pre-training, our method is superior to the previous state-of-the-art text detectors on the arbitrary-shaped text datasets. Codes will be public. Ruijin Liu, Ning Lu 0003, Dapeng Chen, Cheng Li 0040, Zejian Yuan, Wei Peng 0011 |
ACM Multimedia | 3 |
| 2022 | Learning to Predict 3D Lane Shape and Camera Pose from a Single Image via Geometry ConstraintsabstractDetecting 3D lanes from the camera is a rising problem for autonomous vehicles. In this task, the correct camera pose is the key to generating accurate lanes, which can transform an image from perspective-view to the top-view. With this transformation, we can get rid of the perspective effects so that 3D lanes would look similar and can accurately be fitted by low-order polynomials. However, mainstream 3D lane detectors rely on perfect camera poses provided by other sensors, which is expensive and encounters multi-sensor calibration issues. To overcome this problem, we propose to predict 3D lanes by estimating camera pose from a single image with a two-stage framework. The first stage aims at the camera pose task from perspective-view images. To improve pose estimation, we introduce an auxiliary 3D lane task and geometry constraints to benefit from multi-task learning, which enhances consistencies between 3D and 2D, as well as compatibility in the above two tasks. The second stage targets the 3D lane task. It uses previously estimated pose to generate top-view images containing distance-invariant lane appearances for predicting accurate 3D lanes. Experiments demonstrate that, without ground truth camera pose, our method outperforms the state-of-the-art perfect-camera-pose-based methods and has the fewest parameters and computations. Codes are available at https://github.com/liuruijin17/CLGo. Ruijin Liu, Dapeng Chen, Zhiliang Xiong, Zejian Yuan |
AAAI | 2 |
| 2022 | I Can Find You! Boundary-Guided Separated Attention Network for Camouflaged Object DetectionabstractCan you find me? By simulating how humans to discover the so-called 'perfectly'-camouflaged object, we present a novel boundary-guided separated attention network (call BSA-Net). Beyond the existing camouflaged object detection (COD) wisdom, BSA-Net utilizes two-stream separated attention modules to highlight the separator (or say the camouflaged object's boundary) between an image's background and foreground: the reverse attention stream helps erase the camouflaged object's interior to focus on the background, while the normal attention stream recovers the interior and thus pay more attention to the foreground; and both streams are followed by a boundary guider module and combined to strengthen the understanding of boundary. The core design of such separated attention is motivated by the COD procedure of humans: find the subtle difference between the foreground and background to delineate the boundary of a camouflaged object, then the boundary can help further enhance the COD accuracy. We validate on three benchmark datasets that the proposed BSA-Net is very beneficial to detect camouflaged objects with the blurred boundaries and similar colors/patterns with their backgrounds. Extensive results exhibit very clear COD improvements on our BSA-Net over sixteen SOTAs. Peng Li 0064, Haoran Xie 0001, Xuefeng Yan 0001, Dong Liang 0008, Dapeng Chen, Mingqiang Wei, Harry Qin |
AAAI | 6 |
| 2022 | Cross-Modal Retrieval with Heterogeneous Graph EmbeddingabstractConventional methods address the cross-modal retrieval problem by projecting the multi-modal data into a shared representation space. Such a strategy will inevitably lose the modality-specific information, leading to decreased retrieval accuracy. In this paper, we propose heterogeneous graph embeddings to preserve more abundant cross-modal information. The embedding from one modality will be compensated with the aggregated embeddings from the other modality. In particular, a self-denoising tree search is designed to reduce the "label noise" problem, making the heterogeneous neighborhood more semantically relevant. The dual-path aggregation tackles the "modality imbalance" problem, giving each sample comprehensive dual-modality information. The final heterogeneous graph embedding is obtained by feeding the aggregated dual-modality features to the cross-modal self-attention module. Experiments conducted on cross-modality person re-identification and image-text retrieval task validate the superiority and generality of the proposed method. Dapeng Chen, Lin Wu 0001, Harry Qin, Wei Peng 0011 |
ACM Multimedia | 1 |
| 2022 | FNeVR: Neural Volume Rendering for Face AnimationabstractFace animation, one of the hottest topics in computer vision, has achieved a promising performance with the help of generative models. However, it remains a critical challenge to generate identity preserving and photo-realistic images due to the sophisticated motion deformation and complex facial detail modeling. To address these problems, we propose a Face Neural Volume Rendering (FNeVR) network to fully explore the potential of 2D motion warping and 3D volume rendering in a unified framework. In FNeVR, we design a 3D Face Volume Rendering (FVR) module to enhance the facial details for image rendering. Specifically, we first extract 3D information with a well designed architecture, and then introduce an orthogonal adaptive ray-sampling module for efficient rendering. We also design a lightweight pose editor, enabling FNeVR to edit the facial pose in a simple yet effective way. Extensive experiments show that our FNeVR obtains the best overall quality and performance on widely used talking-head benchmarks. Bohan Zeng, Hong Li 0016, Xuhui Liu, Jianzhuang Liu, Dapeng Chen, Wei Peng 0011, Baochang Zhang 0001 |
NeurIPS | 6 |
| 2022 | Pseudo-Pair Based Self-Similarity Learning for Unsupervised Person Re-IdentificationabstractPerson re-identification (re-ID) is of great importance to video surveillance systems by estimating the similarity between a pair of cross-camera person shorts. Current methods for estimating such similarity require a large number of labeled samples for supervised training. In this paper, we present a pseudo-pair based self-similarity learning approach for unsupervised person re-ID without human annotations. Unlike conventional unsupervised re-ID methods that use pseudo labels based on global clustering, we construct patch surrogate classes as initial supervision, and propose to assign pseudo labels to images through the pairwise gradient-guided similarity separation. This can cluster images in pseudo pairs, and the pseudos can be updated during training. Based on pseudo pairs, we propose to improve the generalization of similarity function via a novel self-similarity learning:it learns local discriminative features from individual images via intra-similarity, and discovers the patch correspondence across images via inter-similarity. The intra-similarity learning is based on channel attention to detect diverse local features from an image. The inter-similarity learning employs a deformable convolution with a non-local block to align patches for cross-image similarity. Experimental results on several re-ID benchmark datasets demonstrate the superiority of the proposed method over the state-of-the-arts. Lin Wu 0001, Deyin Liu, Dapeng Chen, ZongYuan Ge, Farid Boussaïd, Mohammed Bennamoun, Jialie Shen 0001 |
IEEE Trans. Image Process. | 4 |
| 2022 | Adaptive Finite-Time Control Scheme for Teleoperation With Time-Varying Delay and UncertaintiesabstractThe communication time delay and uncertain models of robotic manipulators are the major problem in the teleoperation system, which can reduce the performance and stability of the system. This article proposed a novel finite-time adaptive control scheme for position and force tracking performances of the teleoperation system. First, a combined auxiliary error system with position and force tracking errors is designed. Second, a velocity feedback filter is introduced, and a new auxiliary variable function with finite-time structure is designed for controller design. The radial basis function neural network (RBFNN) is applied to estimate the uncertain parts. Then the finite-time adaptive control scheme and adaptive laws are given. Third, based on the Lyapunov method, stability and finite-time performance are demonstrated. And finally, the simulation and experimental studies (with Phantom Ommi devices) are performed and demonstrate the effectiveness of the proposed control scheme on teleoperation position/force tracking. Aiguo Song, Dapeng Chen, Liqiang Fan |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2021 | SSN3D: Self-Separated Network to Align Parts for 3D Convolution in Video Person Re-IdentificationabstractTemporal appearance misalignment is a crucial problem in video person re-identification. The same part of person (e.g. head or hand) appearing on different locations in video sequence weakens its discriminative ability, especially when we apply standard temporal aggregation such as 3D convolution or LSTM. To address this issue, we propose Self-Separated network (SSN) to seek out the same parts in different images. As the name implies, SSN, if trained in an unsupervised strategy, guarantees the selected parts distinct. With a few samples of labeled parts to guide SSN training, this semi-supervised trained SSN seeks out the parts that are human-understandable within a frame and stable across a video snippet. Given the distinct and stable person parts, rather than performing aggregation on features, we then apply 3D convolution across different frames for person re-identification. This SSN + 3D pipeline, dubbed SSN3D, is proved to be efficient through extensive experiments on both synthetic and real data. Xiaoke Jiang, Qichen Li, Wanrong Zheng, Dapeng Chen |
AAAI | 6 |
| 2021 | Gradient Regularized Contrastive Learning for Continual Domain AdaptationabstractHuman beings can quickly adapt to environmental changes by leveraging learning experience. However, adapting deep neural networks to dynamic environments by machine learning algorithms remains a challenge. To better understand this issue, we study the problem of continual domain adaptation, where the model is presented with a labelled source domain and a sequence of unlabelled target domains. The obstacles in this problem are both domain shift and catastrophic forgetting. We propose Gradient Regularized Contrastive Learning (GRCL) to solve the obstacles. At the core of our method, gradient regularization plays two key roles: (1) enforcing the gradient not to harm the discriminative ability of source features which can, in turn, benefit the adaptation ability of the model to target domains; (2) constraining the gradient not to increase the classification loss on old target domains, which enables the model to preserve the performance on old target domains when adapting to an in-coming target domain. Experiments on Digits, DomainNet and Office-Caltech benchmarks demonstrate the strong performance of our approach when compared to the state-of-the-art. Shixiang Tang, Dapeng Chen, Wanli Ouyang |
AAAI | 3 |
| 2021 | Mutual CRF-GNN for Few-Shot LearningabstractGraph-neural-networks (GNN) is a rising trend for fewshot learning. A critical component in GNN is the affinity. Typically, affinity in GNN is mainly computed in the feature space, e.g., pairwise features, and does not take fully advantage of semantic labels associated to these features. In this paper, we propose a novel Mutual CRF-GNN (MCGN). In this MCGN, the labels and features of support data are used by the CRF for inferring GNN affinities in a principled and probabilistic way. Specifically, we construct a Conditional Random Field (CRF) conditioned on labels and features of support data to infer a affinity in the label space. Such affinity is fed to the GNN as the node-wise affinity. GNN and CRF mutually contributes to each other in MCGN. For GNN, CRF provides valuable affinity information. For CRF, GNN provides better features for inferring affinity. Experimental results show that our approach outperforms stateof-the-arts on datasets miniImageNet, tieredImageNet, and CIFAR-FS on both 5-way 1-shot and 5-way 5-shot settings. Shixiang Tang, Dapeng Chen, Lei Bai 0001, Kaijian Liu, Yixiao Ge, Wanli Ouyang |
CVPR | 2 |
| 2021 | Layerwise Optimization by Gradient Decomposition for Continual LearningabstractDeep neural networks achieve state-of-the-art and sometimes super-human performance across various domains. However, when learning tasks sequentially, the networks easily forget the knowledge of previous tasks, known as "catastrophic forgetting". To achieve the consistencies between the old tasks and the new task, one effective solution is to modify the gradient for update. Previous methods enforce independent gradient constraints for different tasks, while we consider these gradients contain complex information, and propose to leverage inter-task information by gradient decomposition. In particular, the gradient of an old task is decomposed into a part shared by all old tasks and a part specific to that task. The gradient for update should be close to the gradient of the new task, consistent with the gradients shared by all old tasks, and orthogonal to the space spanned by the gradients specific to the old tasks. In this way, our approach encourages common knowledge consolidation without impairing the task-specific knowledge. Furthermore, the optimization is performed for the gradients of each layer separately rather than the concatenation of all gradients as in previous works. This effectively avoids the influence of the magnitude variation of the gradients in different layers. Extensive experiments validate the effectiveness of both gradient-decomposed optimization and layer-wise updates. Our proposed method achieves state-of-the-art results on various benchmarks of continual learning. Shixiang Tang, Dapeng Chen, Jinguo Zhu, Shijie Yu, Wanli Ouyang |
CVPR | 2 |
| 2021 | Complementary Relation Contrastive DistillationabstractKnowledge distillation aims to transfer representation ability from a teacher model to a student model. Previous approaches focus on either individual representation distillation or inter-sample similarity preservation. While we argue that the inter-sample relation conveys abundant information and needs to be distilled in a more effective way. In this paper, we propose a novel knowledge distillation method, namely Complementary Relation Contrastive Distillation (CRCD), to transfer the structural knowledge from the teacher to the student. Specifically, we estimate the mutual relation in an anchor-based way and distill the anchor-student relation under the supervision of its corresponding anchor-teacher relation. To make it more robust, mutual relations are modeled by two complementary elements: the feature and its gradient. Furthermore, the low bound of mutual information between the anchor-teacher relation distribution and the anchor-student relation distribution is maximized via relation contrastive loss, which can distill both the sample representation and the inter-sample relations. Experiments on different benchmarks demonstrate the effectiveness of our proposed CRCD. Jinguo Zhu, Shixiang Tang, Dapeng Chen, Shijie Yu, Yakun Liu, Mingzhe Rong, Aijun Yang, Xiaohua Wang 0001 |
CVPR | 3 |
| 2021 | Differentiable Dynamic Wirings for Neural NetworksabstractA standard practice of deploying deep neural networks is to apply the same architecture to all the input instances. However, a fixed architecture may not be suitable for different data with high diversity. To boost the model capacity, existing methods usually employ larger convolutional kernels or deeper network layers, which incurs prohibitive computational costs. In this paper, we address this issue by proposing Differentiable Dynamic Wirings (DDW), which learns the instance-aware connectivity that creates different wiring patterns for different instances. 1) Specifically, the network is initialized as a complete directed acyclic graph, where the nodes represent convolutional blocks and the edges represent the connection paths. 2) We generate edge weights by a learnable module, Router, and select the edges whose weights are larger than a threshold, to adjust the connectivity of the neural network structure. 3) Instead of using the same path of the network, DDW aggregates features dynamically in each node, which allows the network to have more representation power.To facilitate effective training, we further represent the network connectivity of each sample as an adjacency matrix. The matrix is updated to aggregate features in the forward pass, cached in the memory, and used for gradient computing in the backward pass. We validate the effectiveness of our approach with several mainstream architectures, including MobileNetV2, ResNet, ResNeXt, and RegNet. Extensive experiments are performed on ImageNet classification and COCO object detection, which demonstrates the effectiveness and generalization ability of our approach. Quanquan Li, Shaopeng Guo, Dapeng Chen, Aojun Zhou, Fengwei Yu, Ziwei Liu 0002 |
ICCV | 4 |
| 2021 | Online Pseudo Label Generation by Hierarchical Cluster Dynamics for Adaptive Person Re-identificationabstractAdaptive person re-identification (adaptive ReID) targets at transferring learned knowledge from the labeled source domain to the unlabeled target domain. Pseudo-label-based methods that alternatively generate pseudo labels and optimize the training model have demonstrated great effectiveness in this field. However, the generated pseudo labels are inaccurate and cannot reflect the true semantic meaning of the unlabeled samples. We consider such inaccuracy stems from both the lagged update of the pseudo labels as well as the simple criterion of the employed clustering method. To tackle the problem, we propose an online pseudo label generation by hierarchical cluster dynamics for adaptive ReID. In particular, hierarchical label banks are constructed for all the samples in the dataset, and we update the pseudo labels of the sample in each coming mini-batch, performing the model optimization and the label generation simultaneously. A new hierarchical cluster dynamics is built for the label update, where cluster merge and cluster split are driven by a possibility computed by the label propagation. Our method can achieve better pseudo labels and higher reid accuracy. Extensive experiments on Market-to-Duke, Duke-to-Market, MSMT-to-Market, MSMT-to-Duke, Market-to-MSMT, and Duke-to-MSMT verify the effectiveness of our proposed method. Shixiang Tang, Guolong Teng, Yixiao Ge, Kaijian Liu, Harry Qin, Donglian Qi, Dapeng Chen |
ICCV | 8 |
| 2021 | Continual Representation Learning for Biometric IdentificationabstractWith the explosion of digital data in recent years, continuously learning new tasks from a stream of data without forgetting previously acquired knowledge has become increasingly important. In this paper, we propose a new continual learning (CL) setting, namely "continual representation learning", which focuses on learning better representation in a continuous way. We also provide two large-scale multi-step benchmarks for biometric identification, where the visual appearance of different classes are highly relevant. In contrast to requiring the model to recognize more learned classes, we aim to learn feature representation that can be better generalized to not only previously unseen images but also unseen classes/identities. For the new setting, we propose a novel approach that performs the knowledge distillation over a large number of identities by applying the neighbourhood selection and consistency relaxation strategies to improve scalability and flexibility of the continual learning model. We demonstrate that existing CL methods can improve the representation in the new setting, and our method achieves better results than the competitors. Bo Zhao 0038, Shixiang Tang, Dapeng Chen, Hakan Bilen, Rui Zhao 0001 |
WACV | 3 |
| 2021 | Force Display and Tactile Display of Color Image TextureabstractIn haptic interaction technology, texture haptic display is an important part. In order to perceive color image texture better, the haptic display methods of color texture based on force feedback and tactile feedback are proposed in this work. On the one hand, through the study of the physiological and psychological perception characteristics of color information, a new force rendering method of color image texture based on force feedback device is presented. The experimental results of color texture force perception show that the color texture force rendering algorithm in this paper works well. On the other hand, a color texture vibration tactile model is established based on the designed vibrotactile device. Its effectiveness is verified by vibration tactile perception experiment. Finally, for the color texture image of the real object surface, the force and vibrotactile display methods of color texture in this paper are used to conduct perception experiments and compared. Lei Tian 0008, Dapeng Chen, Xiulan Wen, Aiguo Song |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2021 | Multi-Mode Haptic Display of Image Based on Force and Vibration Tactile Feedback IntegrationabstractIn order to enhance the sense of reality haptic display based on image, it is widely expected to express various characteristics of the objects in the image using different kinds of haptic feedback. To this end, a multi-mode haptic display method of image was proposed in this paper, including the multi-feature extraction of image and the image expression with various types of haptic rendering. First, the device structure integrating force and vibrotactile feedbacks was designed for multi-mode haptic display. Meanwhile, the three-dimensional geometric shape, detail texture and outline of the object in the image were extracted by various image processing algorithms. Then, a rendering method for the object in the image was proposed based on the psychophysical experiments on the piezoelectric ceramic actuator. The 3D geometric shape, detail texture and outline of the object were rendered by force and vibration tactile feedbacks, respectively. Finally, these three features of the image were haptic expressed simultaneously by the integrated device. Haptic perception experiment results show that the multi-mode haptic display method can effectively improve the authenticity of haptic perception. Lei Tian 0008, Aiguo Song, Dapeng Chen |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2021 | Person Re-Identification With Deep Kronecker-Product Matching and Group-Shuffling Random WalkabstractPerson re-identification (re-ID) aims to robustly measure visual affinities between person images. It has wide applications in intelligent surveillance by associating same persons' images across multiple cameras. It is generally treated as an image retrieval problem: given a probe person image, the affinities between the probe image and gallery images (P2G affinities) are used to rank the retrieved gallery images. There exist two main challenges for effectively solving this problem. 1) Person images usually show significant variations because of different person poses and viewing angles. The spatial layouts and correspondences between person images are therefore vital information for tackling this problem. State-of-the-art methods either ignore such spatial variation or utilize extra pose information for handling the challenge. 2) Most existing person re-ID methods rank gallery images considering only P2G affinities but ignore the affinities between the gallery images (G2G affinity). Such affinities could provide important clues for accurate gallery image ranking but were only utilized in post-processing stages by current methods. In this article, we propose a unified end-to-end deep learning framework to tackle the two challenges. For handling viewpoint and pose variations between compared person images, we propose a novel Kronecker Product Matching operation to match and warp feature maps of different persons. Comparing warped feature maps results in more accurate P2G affinities. To fully utilize all available P2G and G2G affinities for accurately ranking gallery person images, a novel group-shuffling random walk operation is proposed. Both Kronecker Product Matching and Group-shuffling Random Walk operations are end-to-end trainable and are shown to improve the learned visual features if integrated in the deep learning framework. The proposed approach outperforms state-of-the-art methods on Market-1501, CUHK03 and DukeMTMC datasets, which demonstrates the effectiveness and generalization ability of our proposed approach. Code is available at https://github.com/YantaoShen/kpm_rw_person_reid. Yantao Shen 0002, Tong Xiao 0003, Shuai Yi, Dapeng Chen, Xiaogang Wang 0001, Hongsheng Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2020 | Density-Aware Feature Embedding for Face ClusteringabstractClustering has many applications in research and industry. However, traditional clustering methods, such as K-means, DBSCAN and HAC, impose oversimplifying assumptions and thus are not well-suited to face clustering. To adapt to the distribution of realistic problems, a natural approach is to use Graph Convolutional Networks (GCNs) to enhance features for clustering. However, GCNs can only utilize local information, which ignores the overall characterisitcs of the clusters. In this paper, we propose a Density-Aware Feature Embedding Network (DA-Net) for the task of face clustering, which utilizes both local and non-local information, to learn a robust feature embedding. Specifically, DA-Net uses GCNs to aggregate features locally, and then incorporates non-local information using a density chain, which is a chain of faces from low density to high density. This density chain exploits the non-uniform distribution of face images in the dataset. Then, an LSTM takes the density chain as input to generate the final feature embedding. Once this embedding is generated, traditional clustering methods, such as density-based clustering, can be used to obtain the final clustering results. Extensive experiments verify the effectiveness of the proposed feature embedding method, which can achieve state-of-the-art performance on public benchmarks. Senhui Guo, Dapeng Chen, Xiaogang Wang 0001, Rui Zhao 0001 |
CVPR | 3 |
| 2020 | Learning to Cluster Faces via Confidence and Connectivity EstimationabstractFace clustering is an essential tool for exploiting the unlabeled face data, and has a wide range of applications including face annotation and retrieval. Recent works show that supervised clustering can result in noticeable performance gain. However, they usually involve heuristic steps and require numerous overlapped subgraphs, severely restricting their accuracy and efficiency. In this paper, we propose a fully learnable clustering framework without requiring a large number of overlapped subgraphs. Instead, we transform the clustering problem into two sub-problems. Specifically, two graph convolutional networks, named GCN-V and GCN-E, are designed to estimate the confidence of vertices and the connectivity of edges, respectively. With the vertex confidence and edge connectivity, we can naturally organize more relevant vertices on the affinity graph and group them into clusters. Experiments on two large-scale benchmarks show that our method significantly improves clustering accuracy and thus performance of the recognition models trained on top, yet it is an order of magnitude more efficient than existing supervised methods. Lei Yang 0045, Dapeng Chen, Xiaohang Zhan, Rui Zhao 0001, Chen Change Loy, Dahua Lin |
CVPR | 2 |
| 2020 | COCAS: A Large-Scale Clothes Changing Person Dataset for Re-IdentificationabstractRecent years have witnessed great progress in person re-identification (re-id). Several academic benchmarks such as Market1501, CUHK03 and DukeMTMC play important roles to promote the re-id research. To our best knowledge, all the existing benchmarks assume the same person will have the same clothes. While in real-world scenarios, it is very often for a person to change clothes. To address the clothes changing person re-id problem, we construct a novel large-scale re-id benchmark named Clothes Changing Person Set (COCAS), which provides multiple images of the same identity with different clothes. COCAS totally contains 62,382 body images from 5,266 persons. Based on COCAS, we introduce a new person re-id setting for clothes changing problem, where the query includes both a clothes template and a person image taking another clothes. Moreover, we propose a two-branch network named Biometric-Clothes Network (BC-Net) which can effectively integrate biometric and clothes feature for re-id under our setting. Experiments show that it is feasible for clothes changing re-id with clothes templates. Shijie Yu, Shihua Li 0006, Dapeng Chen, Rui Zhao 0001, Yu Qiao 0001 |
CVPR | 3 |
| 2020 | Adapting Object Detectors with Conditional Domain Normalization
Kun Wang 0056, Xingyu Zeng, Shixiang Tang, Dapeng Chen, Di Qiu, Xiaogang Wang 0001 |
ECCV (11) | 5 |
| 2020 | A Built-In Self-Test Method For MEMS Piezoresistive SensorabstractNowadays, MEMS testing has become a growing problem because it usually needs specific and sophisticated testing equipment and is very time-consuming. To solve this problem, this paper proposes a Built-In Self-Test (BIST) method for membrane MEMS piezoresistive sensor. With the proposed method, an on-chip electric signal can be used as the test stimuli, and process defects of piezoresistive sensor can be diagnosed by analyzing the output response of piezoresistive sensor on chip. The simulation shows that the proposed MEMS BIST scheme can effectively replace the physical testing stimuli with electric signal, thus reduce the dependence on external signal sources and the cost of manufacturing devices. Manhong Zhu, Jia Li 0022, Weibing Wang, Dapeng Chen |
ETS | 4 |
| 2020 | Mutual Mean-Teaching: Pseudo Label Refinery for Unsupervised Domain Adaptation on Person Re-identification
Yixiao Ge, Dapeng Chen, Hongsheng Li 0001 |
ICLR | 2 |
| 2020 | Real-time Continuous Hand Motion Myoelectric Decoding by Automated Data Labeling*abstractIn this paper an automated data labeling (ADL) neural network is proposed to streamline dataset collecting for real-time predicting the continuous motion of hand and wrist, these gestures are only decoded from a surface electromyography (sEMG) array of eight channels. Unlike collecting both the bio-signals and hand motion signals as samples and labels in supervised learning, this algorithm only collects unlabeled sEMG into an unsupervised neural network, in which the hand motion labels are auto-generated. The coefficient of determination (R2) for three DOFs, i.e. wrist flex/extension, wrist pro/supination, hand open/close, was 0.86, 0.89 and 0.87 respectively. The comparison between real motion labels and auto-generated labels shows that the latter has earlier response than former. The results of Fitts’ law test indicate that ADL has capability of controlling multi-DOFs simultaneously even though the training set only contains sEMG data from single DOF gesture. Moreover, no more hand motion measurement needed which greatly helps upper limb amputee imagine the gesture of residual limb to control a dexterous prosthesis. Xuhui Hu, Hong Zeng 0001, Dapeng Chen, Jiahang Zhu, Aiguo Song |
ICRA | 3 |
| 2020 | Deep Transfer Learning Enable End-to-End Steering Angles Prediction for Self-driving CarabstractAutonomous driving has developed rapidly over the last few years. Predicting the steering angle for self-driving car according to different road conditions is very important. There are some endeavors for this topic, including lane detection, object detection on roads, 3-D reconstruction etc., but in our work we focus on a vision based model that directly maps raw input images to steering angles using deep networks and this model don't depend on specifying the features to learn. In this paper, we propose an end-to-end steering angle prediction model based on deep transfer learning and it can accurately predicts steering angles based on input image sequences which are from onboard camera. This prediction model combine two deep learning models including the convolution neural network (CNN) and the long short-term memory (LSTM). The CNN model we use is VGG16 which is based on transfer learning techniques, and pre-trained on Imagenet with good performance. This network is used to extract spatial features of the input image sequences. And the LSTM network is used to capture the temporal information of the provided images. The model we proposed fully considers spatial-temporal information, and fit the nonlinear relationship well between the input images and the steering angles. In order to validate the proposed model, the experimental study is conducted using the real-world dataset which is provided by Udacity. Experimental results show that the proposed model in this paper can efficiently predict the steering angles and clone humans' driving behaviors, and our model has a better performance, higher accuracy, and less training time. Huatao Jiang, Qing Li 0043, Dapeng Chen |
IV | 4 |
| 2020 | Self-paced Contrastive Learning with Hybrid Memory for Domain Adaptive Object Re-IDabstractDomain adaptive object re-ID aims to transfer the learned knowledge from the labeled source domain to the unlabeled target domain to tackle the open-class re-identification problems. Although state-of-the-art pseudo-label-based methods have achieved great success, they did not make full use of all valuable information because of the domain gap and unsatisfying clustering performance. To solve these problems, we propose a novel self-paced contrastive learning framework with hybrid memory. The hybrid memory dynamically generates source-domain class-level, target-domain cluster-level and un-clustered instance-level supervisory signals for learning feature representations. Different from the conventional contrastive learning strategy, the proposed framework jointly distinguishes source-domain classes, and target-domain clusters and un-clustered instances. Most importantly, the proposed self-paced method gradually creates more reliable clusters to refine the hybrid memory and learning targets, and is shown to be the key to our outstanding performance. Our method outperforms state-of-the-arts on multiple domain adaptation tasks of object re-ID and even boosts the performance on the source domain without any extra annotations. Our generalized version on unsupervised object re-ID surpasses state-of-the-art algorithms by considerable 16.7% and 7.9% on Market-1501 and MSMT17 benchmarks. Yixiao Ge, Feng Zhu 0006, Dapeng Chen, Rui Zhao 0001, Hongsheng Li 0001 |
NeurIPS | 3 |
| 2020 | Feel the inside: A haptic interface for navigating stress distribution inside objects
Lifeng Zhu, Rubin Ren, Dapeng Chen, Aiguo Song, Jia Liu 0034, Yin Yang 0004 |
Vis. Comput. | 3 |
| 2019 | Learning to Cluster Faces on an Affinity GraphabstractFace recognition sees remarkable progress in recent years, and its performance has reached a very high level. Taking it to a next level requires substantially larger data, which would involve prohibitive annotation cost. Hence, exploiting unlabeled data becomes an appealing alternative. Recent works have shown that clustering unlabeled faces is a promising approach, often leading to notable performance gains. Yet, how to effectively cluster, especially on a large-scale (i.e. million-level or above) dataset, remains an open question. A key challenge lies in the complex variations of cluster patterns, which make it difficult for conventional clustering methods to meet the needed accuracy. This work explores a novel approach, namely, learning to cluster instead of relying on hand-crafted criteria. Specifically, we propose a framework based on graph convolutional network, which combines a detection and a segmentation module to pinpoint face clusters. Experiments show that our method yields significantly more accurate face clusters, which, as a result, also lead to further performance gain in face recognition. Lei Yang 0045, Xiaohang Zhan, Dapeng Chen, Chen Change Loy, Dahua Lin |
CVPR | 3 |
| 2019 | Memory-Based Neighbourhood Embedding for Visual RecognitionabstractLearning discriminative image feature embeddings is of great importance to visual recognition. To achieve better feature embeddings, most current methods focus on designing different network structures or loss functions, and the estimated feature embeddings are usually only related to the input images. In this paper, we propose Memory-based Neighbourhood Embedding (MNE) to enhance a general CNN feature by considering its neighbourhood. The method aims to solve two critical problems, i.e., how to acquire more relevant neighbours in the network training and how to aggregate the neighbourhood information for a more discriminative embedding. We first augment an episodic memory module into the network, which can provide more relevant neighbours for both training and testing. Then the neighbours are organized in a tree graph with the target instance as the root node. The neighbourhood information is gradually aggregated to the root node in a bottom-up manner, and aggregation weights are supervised by the class relationships between the nodes. We apply MNE on image search and few shot learning tasks. Extensive ablation studies demonstrate the effectiveness of each component, and our method significantly outperforms the state-of-the-art approaches. Suichan Li, Dapeng Chen, Bin Liu 0016, Nenghai Yu, Rui Zhao 0001 |
ICCV | 2 |
| 2019 | Cable-Driven 4-DOF Upper Limb Rehabilitation RobotabstractThis paper developed a 4-degree-of-freedom cable-driven upper limb rehabilitation robot and proposed a control algorithm of the passive training for this robot. Comparing with the conventional cable-driven rehabilitation robot, the workspace of this robot is increased by optimizing the distribution of the cable attachment points and by improving the mechanical design. The rotation structure of the upper arm module can change the distribution of the attachment points as needed, by which the cable tension planner can be satisfied in almost all cases. At the meantime, the internal/external rotation of shoulder joint can be achieved without the change of the cables configuration, which is also important for increasing the workspace and comfortability of utilization. The activities of daily living (ADLs) training can be achieved well without any manual adjustment. The related controller for passive training is designed, which includes a higher controller for trajectory tracking and a lower controller for keeping cable tension as the output of the tension planner in real-time. The passive training experiments are conducted on five healthy subjects of different body size. The results demonstrated that the passive training can be achieved well on different subjects and the cable tension controller is also working effectively. Ke Shi 0006, Aiguo Song, Ye Li 0030, Dapeng Chen |
IROS | 4 |
| 2018 | Group Consistent Similarity Learning via Deep CRF for Person Re-IdentificationabstractPerson re-identification benefits greatly from deep neural networks (DNN) to learn accurate similarity metrics and robust feature embeddings. However, most of the current methods impose only local constraints for similarity learning. In this paper, we incorporate constraints on large image groups by combining the CRF with deep neural networks. The proposed method aims to learn the "local similarity" metrics for image pairs while taking into account the dependencies from all the images in a group, forming "group similarities". Our method involves multiple images to model the relationships among the local and global similarities in a unified CRF during training, while combines multi-scale local similarities as the predicted similarity in testing. We adopt an approximate inference scheme for estimating the group similarity, enabling end-to-end training. Extensive experiments demonstrate the effectiveness of our model that combines DNN and CRF for learning robust multi-scale local similarities. The overall results outperform those by state-of-the-arts with considerable margins on three widely-used benchmarks. Dapeng Chen, Dan Xu 0002, Hongsheng Li 0001, Nicu Sebe, Xiaogang Wang 0001 |
CVPR | 1 |
| 2018 | Video Person Re-Identification With Competitive Snippet-Similarity Aggregation and Co-Attentive Snippet EmbeddingabstractIn this paper, we address video-based person re-identification with competitive snippet-similarity aggregation and co-attentive snippet embedding. Our approach divides long person sequences into multiple short video snippets and aggregates the top-ranked snippet similarities for sequence-similarity estimation. With this strategy, the intra-person visual variation of each sample could be minimized for similarity estimation, while the diverse appearance and temporal information are maintained. The snippet similarities are estimated by a deep neural network with a novel temporal co-attention for snippet embedding. The attention weights are obtained based on a query feature, which is learned from the whole probe snippet by an LSTM network, making the resulting embeddings less affected by noisy frames. The gallery snippet shares the same query feature with the probe snippet. Thus the embedding of gallery snippet can present more relevant features to compare with the probe snippet, yielding more accurate snippet similarity. Extensive ablation studies verify the effectiveness of competitive snippet-similarity aggregation as well as the temporal co-attentive embedding. Our method significantly outperforms the current state-of-the-art approaches on multiple datasets. Dapeng Chen, Hongsheng Li 0001, Tong Xiao 0003, Shuai Yi, Xiaogang Wang 0001 |
CVPR | 1 |
| 2018 | Deep Group-Shuffling Random Walk for Person Re-IdentificationabstractPerson re-identification aims at finding a person of interest in an image gallery by comparing the probe image of this person with all the gallery images. It is generally treated as a retrieval problem, where the affinities between the probe image and gallery images (P2G affinities) are used to rank the retrieved gallery images. However, most existing methods only consider P2G affinities but ignore the affinities between all the gallery images (G2G affinity). Some frameworks incorporated G2G affinities into the testing process, which is not end-to-end trainable for deep neural networks. In this paper, we propose a novel group-shuffling random walk network for fully utilizing the affinity information between gallery images in both the training and testing processes. The proposed approach aims at end-to-end refining the P2G affinities based on G2G affinity information with a simple yet effective matrix operation, which can be integrated into deep neural networks. Feature grouping and group shuffle are also proposed to apply rich supervisions for learning better person features. The proposed approach outperforms state-of-the-art methods on the Market-1501, CUHK03, and DukeMTMC datasets by large margins, which demonstrate the effectiveness of our approach. Yantao Shen 0002, Hongsheng Li 0001, Tong Xiao 0003, Shuai Yi, Dapeng Chen, Xiaogang Wang 0001 |
CVPR | 5 |
| 2018 | Improving Deep Visual Representation for Person Re-identification by Global and Local Image-language Association
Dapeng Chen, Hongsheng Li 0001, Xihui Liu, Yantao Shen 0002, Zejian Yuan, Xiaogang Wang 0001 |
ECCV (16) | 1 |
| 2018 | Show, Tell and Discriminate: Image Captioning by Self-retrieval with Partially Labeled Data
Xihui Liu, Hongsheng Li 0001, Dapeng Chen, Xiaogang Wang 0001 |
ECCV (15) | 4 |
| 2018 | Person Re-identification with Deep Similarity-Guided Graph Neural Network
Yantao Shen 0002, Hongsheng Li 0001, Shuai Yi, Dapeng Chen, Xiaogang Wang 0001 |
ECCV (15) | 4 |
| 2018 | Learning Fixation Point Strategy for Object Detection and ClassificationabstractWe propose a novel recurrent attentional structure to localize and recognize objects jointly. The network can learn to extract a sequence of local observations with detailed appearance and rough context, instead of sliding windows or convolutions on the entire image. Meanwhile, those observations are fused to complete detection and classification tasks. On training, we present a hybrid loss function to learn the parameters of the multi-task network end-to-end. Particularly, the combination of stochastic and object-awareness strategy, named SA, can select more abundant context and ensure the last fixation close to the object. In addition, we build a real-world dataset to verify the capacity of our method in detecting the object of interest including those small ones. Our method can predict a precise bounding box on an image, and achieve high speed on large images. Experimental results indicate that the proposed method can mine effective context by several local observations. Moreover, the precision and speed are easily improved by changing the number of recurrent steps. Source code is available at https://github.com/jielyu/RADCN. Jie Lyu 0001, Zejian Yuan, Dapeng Chen |
ICPR | 3 |
| 2018 | Weighted motion averaging for the registration of multi-view range scans
Jihua Zhu, Yaochen Li, Dapeng Chen, Zhongyu Li 0002, Yongqin Zhang |
Multim. Tools Appl. | 4 |
| 2017 | Fast Pedestrian Detection via Random Projection Features with Shape PriorabstractAccurate pedestrian detection with high speed is always of great interests especially for practical application. Detectors usually follow the feature selection paradigm, and need to first construct rich and diverse features. In particular, current state-of-the-arts generate more channels of feature by convolving the basic feature channels with filter banks, which significantly improves accuracy. In this paper, we propose to apply random projection over the basic feature channels, implicitly selecting feature from a much larger feature space. Our method is more efficient than the ones employing filter banks by avoiding the convolution operation. We further impose shape prior to guide the random projection, making the generated feature be more robust to occlusion, pose variation and scale change. Experimental results on Caltech pedestrian dataset demonstrate the accuracy and efficiency of our method. Compared with thestate-of-arts, our method can achieve 5-10× speedup with comparable accuracy. Zejian Yuan, Dapeng Chen, Jie Lyu 0001 |
WACV | 3 |
| 2017 | Exemplar-Guided Similarity Learning on Polynomial Kernel Feature Map for Person Re-identification
Dapeng Chen, Zejian Yuan, Jingdong Wang 0001, Badong Chen, Gang Hua 0001, Nanning Zheng 0001 |
Int. J. Comput. Vis. | 1 |
| 2017 | Image-based haptic display via a novel pen-shaped haptic device on touch screens
Lei Tian 0008, Aiguo Song, Dapeng Chen |
Multim. Tools Appl. | 3 |
| 2017 | Multi-Timescale Collaborative TrackingabstractWe present the multi-timescale collaborative tracker for single object tracking. The tracker simultaneously utilizes different types of "forces", namely attraction, repulsion and support, to take advantage of their complementary strengths. We model the three forces via three components that are learned from the sample sets with different timescales. The long-term descriptive component attracts the target sample, while the medium-term discriminative component repulses the target from the background. They are collaborated in the appearance model to benefit each other. The short-term regressive component combines the votes of the auxiliary samples to predict the target's position, forming the context-aware motion model. The appearance model and the motion model collaboratively determine the target state, and the optimal state is estimated by a novel coarse-to-fine search strategy. We have conducted an extensive set of experiments on the standard 50 video benchmark. The results confirm the effectiveness of each component and their collaboration, outperforming current state-of-the-art methods. Dapeng Chen, Zejian Yuan, Gang Hua 0001, Jingdong Wang 0001, Nanning Zheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2016 | Similarity Learning with Spatial Constraints for Person Re-identificationabstractPose variation remains one of the major factors that adversely affect the accuracy of person re-identification. Such variation is not arbitrary as body parts (e.g. head, torso, legs) have relative stable spatial distribution. Breaking down the variability of global appearance regarding the spatial distribution potentially benefits the person matching. We therefore learn a novel similarity function, which consists of multiple sub-similarity measurements with each taking in charge of a subregion. In particular, we take advantage of the recently proposed polynomial feature map to describe the matching within each subregion, and inject all the feature maps into a unified framework. The framework not only outputs similarity measurements for different regions, but also makes a better consistency among them. Our framework can collaborate local similarities as well as global similarity to exploit their complementary strength. It is flexible to incorporate multiple visual cues to further elevate the performance. In experiments, we analyze the effectiveness of the major components. The results on four datasets show significant and consistent improvements over the state-of-the-art methods. Dapeng Chen, Zejian Yuan, Badong Chen, Nanning Zheng 0001 |
CVPR | 1 |
| 2016 | Coordinating Multiple Disparity Proposals for Stereo ComputationabstractWhile great progress has been made in stereo computation over the last decades, large textureless regions remain challenging. Segment-based methods can tackle this problem properly, but their performances are sensitive to the segmentation results. In this paper, we alleviate the sensitivity by generating multiple proposals on absolute and relative disparities from multi-segmentations. These proposals supply rich descriptions of surface structures. Especially, the relative disparity between distant pixels can encode the large structure, which is critical to handle the large textureless regions. The proposals are coordinated by point-wise competition and pairwise collaboration within a MRF model. During inference, a dynamic programming is performed in different directions with various step sizes, so the long-range connections are better preserved. In the experiments, we carefully analyzed the effectiveness of the major components. Results on the 2014 Middlebury and KITTI 2015 stereo benchmark show that our method is comparable to state-of-the-art. Dapeng Chen, Yuanliu Liu, Zejian Yuan |
CVPR | 2 |
| 2016 | Haptic Display of Image Based on Multi-Feature ExtractionabstractImage feature extraction is one of the key technologies of image haptic display. In this paper, multi-feature extraction method of the object in image is proposed to improve image-based haptic perception. The multi-feature extraction includes contour shape extraction, pattern extraction and detail texture extraction. Firstly, we use an intrinsic decomposition method to decompose an image into shading image and reflectance image. The reflectance image describes nonillumination affected color patterns spread on the surface. Then, the shading image is utilized in contour shape and detail texture extraction. Contour shape extraction is based on partial differential equation (PDE), to reconstruct three-dimensional (3D) surface model in virtual environments. Detailed texture extraction is based on fractional differential method simultaneously. Finally, the various features extracted above are haptic rendered by different methods. The experimental results show the effectiveness and potentiality of the proposed method for improving the ability of haptic perception and recognition of human in virtual environments. Lei Tian 0008, Aiguo Song, Dapeng Chen, Dejing Ni |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2015 | Similarity learning on an explicit polynomial kernel feature map for person re-identificationabstractIn this paper, we address the person re-identification problem, discovering the correct matches for a probe person image from a set of gallery person images. We follow the learning-to-rank methodology and learn a similarity function to maximize the difference between the similarity scores of matched and unmatched images for a same person. We introduce at least three contributions to person re-identification. First, we present an explicit polynomial kernel feature map, which is capable of characterizing the similarity information of all pairs of patches between two images, called soft-patch-matching, instead of greedily keeping only the best matched patch, and thus more robust. Second, we introduce a mixture of linear similarity functions that is able to discover different soft-patch-matching patterns. Last, we introduce a negative semi-definite regularization over a subset of the weights in the similarity function, which is motivated by the connection between explicit polynomial kernel feature map and the Mahalanobis distance, as well as the sparsity constraint over the parameters to avoid over-fitting. Experimental results over three public benchmarks demonstrate the superiority of our approach. Dapeng Chen, Zejian Yuan, Gang Hua 0001, Nanning Zheng 0001, Jingdong Wang 0001 |
CVPR | 1 |
| 2014 | Description-Discrimination Collaborative Tracking
Dapeng Chen, Zejian Yuan, Gang Hua 0001, Yang Wu 0001, Nanning Zheng 0001 |
ECCV (1) | 1 |
| 2014 | Nanopillar-forest based surface-enhanced Raman scattering substrates
Aida Bao, Haiyang Mao, Jijun Xiong, Zhuojie Chen, Wen Ou, Dapeng Chen |
Sci. China Inf. Sci. | 6 |
| 2014 | A low offset chopper amplifier with three-stage nested Miller configuration
Zhuolei Huang, Weibing Wang, Dapeng Chen |
Sci. China Inf. Sci. | 4 |
| 2013 | Constructing Adaptive Complex Cells for Robust Visual TrackingabstractRepresentation is a fundamental problem in object tracking. Conventional methods track the target by describing its local or global appearance. In this paper we present that, besides the two paradigms, the composition of local region histograms can also provide diverse and important object cues. We use cells to extract local appearance, and construct complex cells to integrate the information from cells. With different spatial arrangements of cells, complex cells can explore various contextual information at multiple scales, which is important to improve the tracking performance. We also develop a novel template-matching algorithm for object tracking, where the template is composed of temporal varying cells and has two layers to capture the target and background appearance respectively. An adaptive weight is associated with each complex cell to cope with occlusion as well as appearance variation. A fusion weight is associated with each complex cell type to preserve the global distinctiveness. Our algorithm is evaluated on 25 challenging sequences, and the results not only confirm the contribution of each component in our tracking system, but also outperform other competing trackers. Dapeng Chen, Zejian Yuan, Yang Wu 0001, Nanning Zheng 0001 |
ICCV | 1 |
| 2012 | Dense Scene Flow Based on Depth and Multi-channel Bilateral Filter
Dapeng Chen, Zejian Yuan, Nanning Zheng 0001 |
ACCV (3) | 2 |
| 2012 | Detecting occlusion boundaries via saliency network
Dapeng Chen, Zejian Yuan, Nanning Zheng 0001 |
ICPR | 1 |
| 2012 | Video object segmentation by clustering region trajectories
Zejian Yuan, Dapeng Chen, Yuehu Liu, Nanning Zheng 0001 |
ICPR | 3 |
| 2002 | Software engineering technology watch
Robert David Cowan, Alan R. McKendall Jr., Ali Mili 0001, Dapeng Chen, V. Janardhana, Terry Spencer |
Inf. Sci. | 6 |