EDBT 2026 Demo / reviewers in the wild / expert
Dihu Chen
dblp:82/10515
· DBLP profile ↗
42ranked-venue papers
0as first author
28since 2021 · last 2026
0000-0001-5432-8149ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 15 since 2021Artificial intelligence and machine learning · 12 · 11 since 2021Systems, architecture and hardware · 11 · 2 since 2021Computer networks · 7 · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PSKNet: Lightweight kernel-aware slice network for real-time stereo depth estimation on edge devices
Bifa Liang, Haifeng Hu 0001, Dihu Chen |
Knowl. Based Syst. | 5 |
| 2026 | SPADNet: lightweight stereo disparity estimation for real-time embedded multimedia systems
Bifa Liang, Yiming Zeng 0008, Haifeng Hu 0001, Dihu Chen |
Multim. Syst. | 4 |
| 2026 | ADA framework for unsupervised domain adaptation person re-identificationabstractDomain shift remains a critical barrier for generalizing person re-identification (ReID) models across datasets. To address this challenge, we present a sparse self-Attention augmented Domain Adaptation (ADA) framework that learns domain-invariant identity features through three key innovations: (1) Sandwich Attention Primitive (SAP), a novel computational unit designed to boost primitive-level domain adaptation. (2) Sparse self-Attention Augmented Bottleneck block (SAAB block), a hierarchical block integrating SAP to enhance adaptation at the architecture level. (3) Scalable Design, if necessary, SAAB block can be flexibly cascaded to construct task-specific ADA framework. Experiments on three benchmarks validate ADA’s superiority: (1) Achieves state-of-the-art performance across domains (e.g., 16.5 % mAP gain on CUHK03 → Market-1501). (2) Demonstrates consistent generalizability and adaptability. Peijun Ye 0002, Dihu Chen |
Pattern Recognit. | 3 |
| 2026 | U-Sparta: A Unified Speech Processing Accelerator for Real-Time ApplicationsabstractWith the advancements in deep learning for speech processing and the rising popularity of smart devices, the demand for edge devices to support multiple speech processing applications is steadily increasing. However, existing unified AI platforms are unsuitable for such edge scenarios, while dedicated accelerators lack flexibility for diverse speech tasks, and current multi-task solutions remain limited in operator support and task coverage. In this work, we propose U-Sparta, a Unified Speech Processing Accelerator for Real-Time Applications. First, a decomposition scheme, comprising three dedicated strategies, is proposed to efficiently support various neural network operators for speech processing. Additionally, an efficient hardware-implemented approach for precise layer-wise fixed-point quantization is introduced, and a range of compression techniques are employed; then we propose hardware-aware expansion to effectively compensate for performance degradation caused by compression. These approaches collectively reduce the model size by an average of 93.1% and computational complexity by 89.9%. The design has been implemented on a Pango PGL50H FPGA, achieving worst case 1.65W and 1.39W on average low power consumption at 50MHz, lower resource utilization compared to single-task related works. It has been further implemented in 55-nm CMOS technology, with a core area of 3.61 mm2, achieving an average power consumption of 22.6 mW and high accuracy across multiple speech processing applications, including voice activity detection (VAD), keyword spotting (KWS), speaker verification (SV), environmental sound classification (ESC), speech enhancement (SE), speech separation (SS), and speech recognition (SR). Haohai Yu, Keyan He, Dihu Chen |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2026 | RCAENet: Residual Convolutional and Attention-Enhanced Stereo Matching for Real-Time Depth Estimation on Edge DevicesabstractAs a core technology in real-time video processing and intelligent surveillance, stereo matching provides essential depth perception capabilities for multimedia applications. However, high-precision stereo networks often come with significant computational costs, making real-time inference on power- and memory-constrained edge devices challenging. On the other hand, lightweight real-time networks still struggle with accuracy limitations. To address this challenge, we propose RCAENet, a high-performance stereo network designed for real-time and high-accuracy depth estimation on edge devices. To enhance feature extraction efficiency, we introduce the Residual Convolutional Feature Extraction (RCFE) module, which replaces conventional convolutional layers to capture more expressive features while maintaining computational efficiency. Additionally, we propose the Enhanced Adaptive Upsampling (EAU) module, which integrates channel and spatial attention mechanisms to improve feature fusion and disparity refinement. Furthermore, we design an Enhanced 3D CNN (E3DC) along with the Cost Aggregation and Residual Attention (CA-ResAgg) module for cost volume regularization. This module incorporates residual aggregation and efficient channel attention to further enhance disparity estimation accuracy. Built upon these components, RCAENet features a multi-scale architecture that effectively balances accuracy and efficiency. Extensive experiments demonstrate that these innovations enable RCAENet to achieve real-time inference on edge devices while maintaining state-of-the-art depth accuracy. Bifa Liang, Haifeng Hu 0001, Dihu Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2025 | MDGFusion: A Mask-Guided Depth-Reliable Generative Approach for 3D Object Detection via 4D Radar-Camera Fusion
Zhijie Zheng 0003, Dihu Chen |
PRICAI (5) | 5 |
| 2025 | UniPT: A Unified Representation Pre-training for Multi-dataset 3D Object Detection
Zhijie Zheng 0003, Kang Lin, Wei Zhou 0042, Zhongyuan Qiu, Yali Zhao, Huan Qin, Dihu Chen |
PRICAI (5) | 10 |
| 2025 | From multi-scale grids to dynamic regions: Dual-relation enhanced transformer for image captioning
Wei Zhou 0042, Chuanle Song, Dihu Chen, Haifeng Hu 0001, Chun Shan |
Knowl. Based Syst. | 3 |
| 2025 | DRTN: Dual Relation Transformer Network with feature erasure and contrastive learning for multi-label image classification
Wei Zhou 0042, Kang Lin, Zhijie Zheng 0003, Dihu Chen, Haifeng Hu 0001 |
Neural Networks | 4 |
| 2025 | Sparse-attention augmented domain adaptation for unsupervised person re-identification
Peijun Ye 0002, Dihu Chen |
Pattern Recognit. Lett. | 4 |
| 2025 | RAFDet: Range View Augmented Fusion Network for Point-Based 3D Object DetectionabstractIn recent years, point-based methods have achieved promising performance on 3D object detection task. Although effective, they still suffer from the inherent sparsity of point cloud, which makes it challenging to distinguish objects with backgrounds only relying on the view of raw point. To this end, we propose a straightforward yet effective multi-view fusion network termed RAFDet to alleviate this issue. The core idea of our method lies in combining the merits of raw point and its range view to enhance the representation learning for sparse point cloud, thus mitigating the sparsity problem and boosting the detection performance. In particular, we introduce a novel bidirectional attentive fusion module to equip sparse point with interacted fine-grained semantic clues during feature learning process. Then, we devise the range-view augmented fusion module to fully exploit the supplementary relationship between different perspectives with the aim of enhancing original point-view features. In the end, a single-stage detection head is utilized to predict final 3D bounding boxes based on the enhanced semantics. We have evaluated our method on the popular KITTI Dataset, DAIR-V2X Dataset and Waymo Open Dataset. Experimental results on the above three datasets demonstrate the effectiveness and robustness of our approach in terms of detection performance and model complexity. Zhijie Zheng 0003, Kang Lin, Haifeng Hu 0001, Dihu Chen |
IEEE Trans. Multim. | 6 |
| 2025 | AMVFNet: Attentive Multi-View Fusion Network for 3D Object DetectionabstractPillar-based method is significant in the field of LiDAR-based 3D object detection which could directly make use of efficient 2D backbone and save computational resources during reference. Existing methods usually sequentially project the original point clouds into the cylindrical view or the bird-eye view for feature extraction. However, the former suffers from obscured problems and the scales of instances vary greatly with distance, while the latter leads to considerable confusion problems due to the loss of semantic information caused by the sparsity of the projected point cloud. In this article, we present a novel and efficient two-stage point-pillar hybrid architecture named Attentive Multi-View Fusion Network (AMVFNet), in which we abstract features from all cylindrical view, bird-eye view, and raw point clouds. Rather than designing more complex modules to solve the problems inherent in the single-view approach, our multi-view fusion architecture effectively combines the strengths of multiple perspectives to improve performance at a more fundamental level. Besides, to compensate for quantization distortion caused by projection operations, we propose attentive feature enhancement layers to further improve the capability of contextual information capturing. Extensive experiments on the KITTI detection benchmark illustrate that our proposed AMVFNet achieves competitive performance compared with other SOTA 3D object detectors. Haifeng Hu 0001, Dihu Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | Temporal and Semantic Correlation Network for Weakly-Supervised Temporal Action LocalizationabstractWeakly-Supervised Temporal Action Localization (WTAL) aims to identify the temporal boundaries and classify actions in untrimmed videos using only video-level labels during training. Despite recent progress, many existing approaches primarily follow a localization-by-classification pipeline, treating snippets as independent instances and thus exploiting only limited contextual information. Besides, these methods struggle to capture multi-scale temporal information and neglect both the internal temporal structures within videos and the semantic consistency between videos, resulting in misclassification and inaccurate localization. To address these limitations, we introduce a novel Temporal and Semantic Correlation Network (TSC-Net) for WTAL task, which can be trained end-to-end. First, we propose a Multi-Scale Features Integration Pyramid (MFIP) module to integrate multi-scale temporal features, effectively addressing the challenge of missed detections caused by short action durations. Furthermore, we design a Temporal Correlation Enhancement (TCE) branch to enhance segment correlations by video-level temporal structures to improve the completeness of action localization. Finally, a Dataset-Wide Semantic Awareness (DSA) branch is designed to construct and propagate a dataset-level action semantics bank, enhancing the model’s awareness of semantic consistency in actions. Extensive experiments show that TSC-Net outperforms most existing WTAL methods, achieving an average mAP of 46.3% on the THUMOS-14 dataset and 26.5% on the ActivityNet1.2 dataset. Detailed ablation studies further confirm the effectiveness of each component in our model. The code and models are publicly available at https://github.com/linkang-els/TSC-Net-main . Kang Lin, Wei Zhou 0042, Zhijie Zheng 0003, Dihu Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | Text-and-Image Learning Transformer for Cross-Modal Person Re-IdentificationabstractText-based person re-identification aims to find the target person from a large pedestrian gallery with the given natural language description. Previous works mainly focus on embedding salient textual and visual representations in a common latent space by utilizing the dual-path structure or parameter-shared network. However, they still lack the ability to effectively extract fine-grained unimodal features as well as fuse the cross-modal data, leading to the increase of misaligned cases. To settle these issues, we propose a text-and-image implicit learning Transformer (TILT) to eliminate textual anisotropy and enhance the cross-modal alignment from both domains based on the bi-direction multi-modal encoders. Specifically, we apply the pre-trained multi-modal embedding module to overcome the unimodal anisotropy problem with contrastive learning, and map fine-grained features with dual encoder in bi-directional masking. Then, we design the cross-modal interaction encoder to comprehensively mine implicit cross-modal relations by reconstructing masked tokens, and fuse rich multi-modal knowledge in a common space. In addition, the cross-modal similarity matching module is proposed to optimize the intra-domain classification and decrease the inter-domain divergence. Extensive experiments are conducted on three public benchmarks CUHK-PEDES, ICFG-PEDES, and RSTPReid to verify the effectiveness of our proposed framework. Results prove that our model outperforms state-of-the-art methods on all metrics. Tinghui Wu, Shuhe Zhang, Dihu Chen, Haifeng Hu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | SHP-FsNTT: A Scalable and High-Performance NTT Accelerator Based on the Four-step AlgorithmabstractLattice-Based Cryptography (LBC) emerges as a powerful cryptographic primitive, offering a solution for post-quantum security. Within LBC schemes, one of the most computationally intensive tasks is polynomial multiplication, which can be accelerated through the Number Theoretic Transform (NTT). This paper proposes SHP-FsNTT, a scalable, dynamically configurable and high-performance hardware accelerator based on four-step NTT algorithm to support both NTT and inverse NTT (INTT). SHP-FsNTT leverages pipeline parallelism and data parallelism, and optimizes the memory access pattern to avoid the implementation of a matrix transposition unit for the four-step algorithm. The proposed design achieves remarkable area-time efficiency improvement compared with state-of-the-art works on FPGA. Weicong Lu, Dihu Chen |
ISCAS | 4 |
| 2024 | A lightweight distillation recurrent convolution network on FPGA for real-time video super-resolution
Zhaowen Zheng, Yuqiao Huang, Dihu Chen |
Multim. Syst. | 3 |
| 2024 | Mining Semantic Information With Dual Relation Graph Network for Multi-Label Image ClassificationabstractThe purpose of multi-label image classification is to assign multiple labels for multiple objects presented in one image. Recent research efforts exploit graph convolution network (GCN) to learn the label co-occurrence dependencies for enhancing the semantic representation. Although these methods have achieved promising results, they can not capture the intrinsic correlation between objects in images and do not consider the inter-channel relationship. In addition, the previous methods treat each single image independently and fail to explore the relationship between different images. To address the above challenges, we propose a novelDualRelationGraphNetwork (DRGN) model, which adopts a double branch structure to excavate rich semantic information from intra-image and cross-image simultaneously. Specifically, we first develop an intra-image channel-relation mining (ICM) module to mine the inter-channel relationship in features while learning the importance of different channels. Secondly, we design a new GCN-based intra-image spatial-relation exploring (ISE) module to capture the correlation between objects in individual image. Notably, ISE module and ICM module can complement and promote each other from the spatial and channel dimensions of images to improve the correlation between objects in individual image. Thirdly, we propose a novel GCN-based cross-image semantic learning (CSL) module to learn the semantic relationship between different images in the mini-batch. Through graph reasoning, our CSL module can iteratively refine input image features by acquiring common semantic information from other images in the mini-batch. Extensive experiments on the MS-COCO 2014, PASCAL VOC 2007, and VG-500 datasets demonstrate that the proposed DRGN model outperforms current state-of-the-art methods. Wei Zhou 0042, Weitao Jiang, Dihu Chen, Haifeng Hu 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Self-Supervised Consistency Based on Joint Learning for Unsupervised Person Re-identificationabstractRecently, unsupervised domain adaptive person re-identification (Re-ID) methods have been extensively studied thanks to not requiring annotations, and they have achieved excellent performance. Most of the existing methods aim to train the Re-ID model for learning a discriminative feature representation. However, they usually only consider training the model to learn a global feature of a pedestrian image, but neglecting the local feature, which restricts further improvement of model performance. To address this problem, two local branches are added to the networks, aiming to allow the model to focus on the local feature containing identity information. Furthermore, we propose a self-supervised consistency constraint to further improve robustness of the model. Specifically, the self-supervised consistency constraint uses the basic data augmentation operations without other auxiliary networks, which can improve performance of the model effectively. Then, a learnable memory matrix is designed to store the mapping vectors that maps person features into probability distributions. Finally, extensive experiments are conducted on multiple commonly used person Re-ID datasets to verify the effectiveness of the proposed generative adversarial networks fusing global and local features. Experimental results reveal that our method achieves results comparable to state-of-the-art methods. Xulei Lou, Tinghui Wu, Haifeng Hu 0001, Dihu Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | PSA-Det3D: Pillar set abstraction for 3D object detection
Zhijie Zheng 0003, Haifeng Hu 0001, Dihu Chen |
Pattern Recognit. Lett. | 6 |
| 2023 | DTSSD: Dual-Channel Transformer-Based Network for Point-Based 3D Object DetectionabstractIn the field of 3D object detection, previous methods mainly utilize one channel feature encoding network to extract point-wise features. Despite the effectiveness, we find that only leveraging one channel encoding network is not sufficient and impedes the detection performance. To this end, we propose a dual-channel transformer-based feature encoding network, which integrates both set abstraction layer and transformer block as backbone. It enables the model to exploit fine-grained as well as long-range contextual information of objects, thus providing complementary relationship of two methods. In addition, a centroid estimation module is introduced to obtain powerful representation of the whole object. Finally, considering the significance of point density, which is crucial for detection performance, we propose a central density-aware enhancement module to equip center features with distinct density features. Experimental results on KITTI dataset show the effectiveness of our proposed method. Zhijie Zheng 0003, Haifeng Hu 0001, Dihu Chen |
IEEE Signal Process. Lett. | 5 |
| 2023 | Diverse Feature Learning Network With Attention Suppression and Part Level Background Suppression for Person Re-IdentificationabstractIn this paper, we propose a Diverse Feature Learning Network with Attention Suppression and Part Level Background Suppression (DFLN) for person re-identification (ReID). DFLN includes two key components: attention suppression mechanism (ASM) and part level background suppression mechanism (PLBSM). Firstly, despite attention mechanism has made great progress in current state-of-the-art ReID methods, they can only pay attention to the most salient region but ignore other discriminative information limiting the diversity of networks, which is not optimal for ReID due to the models tend to match persons by diverse clues (e.g., legs, arms, body, logo of clothes). To tackle the limitation above, we propose the ASM to assist the network to make full use of the most salient features and capture the other sub-salient features, so as to attain diverse features to improve the network performance. Secondly, we adopt a novel PLBSM to develop the part-based method which is proved effective for enhancing the diversity of ReID network. The PLBSM consists of a part feature refined module and a background suppression loss function, and aims to attain pure part level feature by filtering background clutter. Our DFLN integrates the ASM and part-based method developed by PLBSM into an end-to-end network and is able to extract robust diversity feature representations leading to higher performance. Extensive experimental results demonstrate the effectiveness of each component and our method achieves state-of-the-art results on mainstream person re-identification datasets. Shengrong Yang, Weihong Liu, Yangbin Yu, Haifeng Hu 0001, Dihu Chen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Attention-Augmented Memory Network for Image Multi-Label ClassificationabstractThe purpose of image multi-label classification is to predict all the object categories presented in an image. Some recent works exploit graph convolution network to capture the correlation between labels. Although promising results have been reported, these methods cannot learn salient object features in the images and ignore the correlation between channel feature maps. In addition, the current researches only learn the feature information within individual input image, but fail to mine the contextual information of various categories from the dataset to enhance the input feature representation. To address these issues, we propose an A ttention- A ugmented M emory N etwork ( AAMN ) model for the image multi-label classification task. Specifically, we first propose a novel categorical memory module to excavate the contextual information of various categories from the dataset to augment the current input feature. Secondly, we design a new channel-relation exploration module to capture the inter-channel relationship of features, so as to enhance the correlation between objects in the images. Thirdly, we develop a spatial-relation enhancement module to model second-order statistics of features and capture long-range dependencies between pixels in feature maps, so as to learn salient object features. Experimental results on standard benchmarks, including MS-COCO 2014, PASCAL VOC 2007, and VG-500, demonstrate the effectiveness and superiority of AAMN model, which outperforms current state-of-the-art methods. Wei Zhou 0042, Yanke Hou, Dihu Chen, Haifeng Hu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | Attention-Guided Multi-Clue Mining Network for Person Re-identification
Yangbin Yu, Shengrong Yang, Haifeng Hu 0001, Dihu Chen |
Neural Process. Lett. | 4 |
| 2022 | Two-Branch Asymmetric Model With Alternately Clustering for Unsupervised Person Re-IdentificationabstractIn the field of unsupervised person Re-identification (Re-ID), mainstream methods adopt cluster algorithm to generate pseudo labels for training. Despite the effectiveness, the cluster algorithm generates noisy labels, which are retained in further model updating and hinder higher performance. To solve this problem, we propose a Two-branch Asymmetric Model with Alternately Clustering. Specifically, the designed Alternately Clustering (AC) strategy leverages a two-branch model to cluster different pseudo labels for each branch, which prevents the continuous existence of identical noisy labels. To establish a mapping between the two label sets, pseudo label mapping constraint (PLMC) module is devised, which helps retain reliable pseudo labels. Our method improves the quality of generated pseudo labels by keeping noisy labels changing and retaining the reliable ones. Experimental results demonstrate that our proposed method outperforms the state-of-the-art methods. Yangbin Yu, Haifeng Hu 0001, Dihu Chen |
IEEE Signal Process. Lett. | 4 |
| 2021 | Image translation with dual-directional generative adversarial networksabstractAbstract Image‐to‐image translation is a class of vision and graphics problems where the goal is to learn the mapping between input images and output images. However, due to the unstable training and limited training samples, many existing GAN‐based works have difficulty in producing photo‐realistic images. Herein, dual‐directional generative adversarial networks are proposed, which consist of four adversarial networks, to produce images of high perceptual quality. In this framework, self‐reconstruction strategy is used to construct auxiliary sub‐networks, which impose more effective constraints on encoder‐generator pairs. Using this idea, this model can increase the use ratio of paired data conditioned on the same dataset and obtain well‐trained encoder‐generator pairs with the help of the proposed cross‐network skip connections. Moreover, the proposed framework not only produces realistic images but also addresses the problem where condition GAN produces sharp images containing many small, hallucinated objects. Training on multiple supervised datasets, convincing evidences are shown to prove that this model can achieve compelling results by latently learning a common feature representation. Qualitative and quantitative comparisons against other methods, demonstrate the effectiveness and superiority of the method. Congcong Ruan, Liuchun Yuan, Haifeng Hu 0001, Dihu Chen |
IET Comput. Vis. | 4 |
| 2021 | Boundary Adjusted Network Based on Cosine Similarity for Temporal Action Proposal Generation
Jingye Zheng, Dihu Chen, Haifeng Hu 0001 |
Neural Process. Lett. | 2 |
| 2021 | Part-Relation-Aware Feature Fusion Network for Person Re-IdentificationabstractThe research of part-based methods has been proven as an effective way in person re-identification (Re-ID) task. However, in existing part-based Re-ID methods, the informative interactions and potential associations among parts are neglected, which demands further study. To fill this gap, we propose a novel Part-relation-aware Feature Fusion Network (PFFN) which achieves a part-level feature fusion and enhances the discrimination of part features by fully employing helpful information from associations among parts. More specifically, a Dual-stage Attention (DA) module, consisting of spatial and part-based channel attention, is proposed to exploit complementary benefits of two kinds of attention information, thereby facilitating model with learning more discriminative features. Furthermore, Part-relation Exploitation (PE) module is proposed to learn relation-aware part features where correlative information among parts are fully employed, thereby bringing a noticeable improvement in performance. Extensive experiments are conducted on four mainstream Re-ID datasets to verify the superiority of PFFN. Compared with baseline model, PFFN has gained rank-1 accuracy improvement of 18.2% on MSMT17-v2, 12.1% on the CUHK03-Labeled, 5.6% on DukeMTMC-reid and 1.3% on Market1501, compellingly validating its effectiveness. Yanke Hou, Sicheng Lian, Haifeng Hu 0001, Dihu Chen |
IEEE Signal Process. Lett. | 4 |
| 2021 | Multiscale Omnibearing Attention Networks for Person Re-IdentificationabstractThe past few years in the fields of Person Re-Identification (RE-ID) have seen attention mechanism receives enormous interest as it has superior performance in obtaining discriminative feature representations. However, a wide range of state-of-the-art RE-ID attention models only focus on one-dimensional attention design method, e.g. spatial attention and channels attention, hence the produced attention maps are neither detailed enough nor discriminative enough to capture complicated interactions of visual parts. Developing multi-scale attention mechanism for RE-ID, an under-studied approach, becomes a practicable method to overcome this deficiency. Toward this goal, we propose a Multiscale Omnibearing Attention Networks (MOAN) for RE-ID which is capable of utilizing the complex fusion information acquired from the multiscale attention mechanism with features being more representative. Specifically, MOAN takes full advantage of multi-sized convolution filters to obtain discriminative holistic and local feature maps, and adaptively conducts feature information augmentation by introducing an Omnibearing Attention (OA) module. Through the OA module, spatial attention and channel attention are integrated together in a unique way where they work in a complementary way. To sum up, MOAN not only inherits the merit of two kinds of attention mechanism but also performs well in extracting comprehensive feature information. Furthermore, taking into account the robustness of model performance, we formulate a Random Drop (RD) Function to facilitate training MOAN and further increase the diversity of training model for adaptation. Furthermore, to achieve end-to-end training, we utilize trainable parameters to take place of initial fixed parameters, and the model performance is experimentally promoted. Extensive experiments have been carried out on the four mainstream RE-ID datasets. As the result shows, our method with re-ranking achieves rank-1 accuracy of 92.29% on CUHK03-NP, 97.45% on Market-1501, 93.81% on DukeMTMC-reID and 81.53% on MSMT17-V2, outperforming the state-of-the-art methods and confirming the effectiveness of our method. Yewen Huang, Sicheng Lian, Haifeng Hu 0001, Dihu Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | W-Band Synthesized Modulator and Demodulator with Wideband Performance in 65-nm CMOSabstractA high image-rejection in-phase/quadrature (IQ) modulator and a wideband demodulator with large dynamic range operating in W-band is implemented in 65-nm CMOS technology. The modulator demonstrates a flat conversion gain of 12.5±0.7dB, a minimum image-rejection ratio of 40 dBc and a maximum LO-to-RF leakage of 38 dB from 90 to 98 GHz. The conversion gain of the demodulator is digitally controlled from 15 to 46 dB with a noise figure from 13.5 to 11.5 dB and an input 1dB compression point (IP1dB) from -13 to -39 dBm. The modulator and the demodulator are synthesized by two on-chip single-pole-double-throw (SPDT) switches and one IF switch. The entire system occupies 1.4 mm2chip area with a total power consumption of 165 mW. Zhenpeng Zheng, Xiangyu Meng 0003, Dihu Chen |
ISCAS | 5 |
| 2020 | Deeply Associative Two-Stage Representations Learning Based on Labels Interval Extension Loss and Group Loss for Person Re-IdentificationabstractPerson Re-identification (ReID) aims to match people across non-overlapping camera views in a public space, which is usually regarded as an image retrieval problem to match query images with pedestrian images in the gallery. It is challenging since many difficulties exist such as pose misalignments, occlusions, similar appearance when detecting people. Existing researches on ReID mainly focus on two major problems: representation learning and metric learning. In this paper, we target at learning discriminative representations and make two contributions in total. (i) We propose a novel architecture named Deeply Associative Two-stage Representations Learning (DATRL). It contains the global re-initialization stage and fully-perceptual classification stage employing two identical CNNs associatively at the same time. On the global stage, we take on the backbone of one deep CNN e.g., dozens of layers in the front of Resnet-50 as a normal re-initialization subnetwork. Meanwhile, we apply our own proposed 3D-transpose technique into the backbone of the other CNN to form the 3D-transpose re-initialization subnetwork. The fully-perceptual stage is actually made up of the leftover layers of the original CNNs. On this stage, we take both the global representations learned at multiple hierarchies and the local representations uniformly-partitioned on the highest conv-layer into consideration, and then optimizing them separately for classification. (ii) We introduce a new joint loss function in which our proposed Labels Interval Extension loss (LIEL) and Group loss (GL) are combined to enhance the performance of gradient decent as well as increasing the distances between image features with different identities. We apply the above DATRL, LIEL and GL to ReID thus obtaining DATRL-ReID. Experimental results on four datasets CUHK03, Market-1501, DukeMTMC-reID and MSMT17-V2 demonstrate that DATRL-ReID shows excellent performance in improving recognition accuracy and is superior to state-of-the-art methods. Yewen Huang, Yi Huang 0035, Haifeng Hu 0001, Dihu Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Three-Dimension Transmissible Attention Network for Person Re-IdentificationabstractIn this work, we propose a Three-Dimensional Transmissible Attention Network (3DTANet) for Person Re-Identification, which can transmit the attention information from layer to layer and attend to the person image from a three-dimensional perspective. Main contributions of the 3DTANet are: (i) A novel Transmissible Attention (TA) mechanism is introduced, which can transfer attention information between convolution layers. Different from traditional attention mechanism, not only can it convey accumulated attention information layer by layer but also guide the network to retain holistic attention information. (ii) We propose a Three-Dimension Attention (3DA) mechanism, which is capable of extracting a three-dimensional attention map. While previous researches on image attention mechanism extracts channel or spatial attention information separately, 3DA mechanism pays attention to channel and spatial information simultaneously, thereby making them play better complementary role in attention extraction. (iii) A new loss function named L2-norm Multi-labels Loss (L2ML) is applied to acquire higher recognition accuracy calculated by multi labels of same ID and corresponding feature representation. Quite different from the common loss functions, L2-norm Multi-labels Loss is specifically good at optimizing feature distance. In brief, 3DTANet gains two-fold benefit toward higher accuracy. For one thing, the attention information is informative and can be transmitted, feature being more representative. For another, our model is computationally lightweight and can be easily applied to real scenarios. We extensively conduct experiments on four Person Re-Identification benchmark datasets. Our model achieves rank-1 accuracy of 87.50% on CUHK03, 96.23% on Market-1501, 92.50% on DukeMTMC-reID and 76.60% on MSMT17-V2 respectively. The results confirm that the 3DTANet can extract more representative features and attain a higher recognition accuracy, outperforming the state-of-the-art methods. Yewen Huang, Sicheng Lian, Suian Zhang, Haifeng Hu 0001, Dihu Chen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2019 | Channel and Constraint Compensation for Generative Adversarial Networks
Wei Wang 0210, Haifeng Hu 0001, Dihu Chen |
PRCV (1) | 3 |
| 2019 | Hierarchical extended collaborative representation based classification for single-sample face recognitionabstractCollaborative representation based classification (CRC) has been widely used and shown good performance in face recognition (FR). Afterwards, hierarchical representation based classification has recently been proposed and aims to enhance the classification performance of the CRC method. However, these methods highly depend on the over‐complete dictionary comprised of sufficient training samples, and cannot be directly applied for single‐sample FR. In this study, the authors propose a novel CRC‐based FR framework to address this issue, which is named hierarchical extended collaborative representation based classification (HECRC). Firstly, they integrate hierarchical representation based model with low‐rank constrained variation dictionaries. Secondly, they select training samples that are the nearest neighbours of test images to obtain a more discriminative training dictionary, where an adaptive scheme is introduced to select proper samples automatically instead of setting a predefined number in traditional methods. Finally, the refined training dictionary and the learned variation dictionaries are jointly utilised to represent the test sample. Moreover, they combined the proposed HECRC with deep features to further improve the recognition rate. Experiments have been conducted on the AR and FERET datasets, and the results show that the proposed method has a substantial improvement over existing algorithms for single‐sample FR. Yuelai Yuan, Dihu Chen, Haifeng Hu 0001, Lingshuang Du |
IET Comput. Vis. | 2 |
| 2018 | A 70-nA 13-ppm/°C All-MOSFET Voltage Reference for Low-Power IoT SystemsabstractThis paper presents a low-power All-MOSFET voltage reference implemented on a 0.18-μm standard CMOS technology. In order to improve the temperature coefficient (TC) of voltage reference, a TC compensation technique based on controlling bulk voltage is proposed. The proposed voltage reference achieves a TC of 13 ppm/°C from -40 °C to 125 °C while dissipating a supply current of 70 nA in normal temperature. The line regulation is 0.02%/V when the supply voltage varies from 1.3 V to 2.1 V, and the power supply rejection ratio (PSRR) at 100 Hz is 74 dB due to the cascode current mirror. Moreover, the current mirror can be reconfigured easily so that the output voltage can be trimmed in this design. Jianping Guo 0004, Siji Huang, Bing Mo, Dihu Chen |
ISCAS | 7 |
| 2018 | Improved Synthesis of Compressor Trees in High-Level Synthesis for Modern FPGAsabstractIn this paper, an approach to synthesize compressor trees in high-level synthesis is proposed. We target the modern field-programmable gate arrays, which integrate carry chains and support fast ternary adders. Two main improvements are achieved in our approach: 1) based on the proposed modified bitmask analysis, we perform bit-level numerical optimizations to shrink the scale of generated compressor trees and achieve a better area-delay performance; 2) by estimating the arrival time of each multi-input addition operand, we combine the use of generalized parallel counters and ternary adders in compressor trees to further reduce the area while maintaining a similar delay performance. A series of experiments shows that our approach reduces the area significantly while maintaining similar delay performance, as compared to the existing approaches. Le Tu, Yuelai Yuan, Kan Huang, Xiaoqiang Zhang 0010, Dihu Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2017 | Improved Synthesis of Compressor Trees on FPGAs in High-Level SynthesisabstractIn this paper, an approach to synthesize compressor trees in High-level Synthesis (HLS) for FPGAs is proposed. Our approach utilizes the bit-level information to improve the compressor tree synthesis. To obtain the bit-level information targeting compressor tree synthesis, a modified bitmask analysis technique based on prior work is proposed. A series of experimental results show that, compared to the existing heuristic, the average reductions of area and delay are 22.96% and 7.05%. The reductions increase to 29.97% and 9.07% respectively, when the carry chains in FPGAs are utilized to implement the compressor trees. Le Tu, Yuelai Yuan, Kan Huang, Xiaoqiang Zhang 0010, Dihu Chen |
FCCM | 6 |
| 2017 | Improved Nauta transconductor for wideband intermediate-frequency gm-C filterabstractAn improved Nauta transconductor, with output resistance for differential-mode output signals insensitive to process and tuning voltage variations, is presented in this paper. Comparing with the classical Nauta transconductor, the proposed transconductor introduces 4 auxiliary inverters only. It can increase the differential-mode output resistance, and keep the common-mode output resistance nearly no changed at the same time. The proposed transconductor has been implemented in a 7th-order transconductance-C (gm-C) band-pass filter (BPF) in a standard 0.18-μm CMOS technology. The proposed filter has achieved a bandwidth of 20 MHz under a 15-MHz center frequency. Monte Carlo simulation results show that with the same tuning voltage range, the gain deviation of the filter with proposed transconductors is reduced 4.4 dB comparing with the counterpart based on classical Nauta transconductors, and the sensitivity of the quality (Q) factor of filter poles to tuning voltage is also significantly reduced. Measurement results show that the pass-band ripple of the proposed filter is reduced 1 dB, and the stop-band attenuation is reduced 8 dB at 30 MHz. Jianghui Deng, Zhuojian Fu, Dihu Chen, Xian Tang, Jianping Guo 0004 |
ISCAS | 4 |
| 2017 | A mechanism for detecting on-chip radio frequency interference of field-programmable gate array
Hengfei Zhong, Zhuoquan Huang, Dihu Chen |
Integr. | 3 |
| 2013 | A gradual scheduling framework for problem size reduction and cross basic block parallelism exploitation in high-level synthesisabstractIn High-level Synthesis, scheduling has a critical impact on the quality of hardware implementation. However, the schedules of different operations are actually having unequal impacts on the Quality of Result. Based on this fact, we propose a novel scheduling framework, which is able to schedule the operations separately according their significance to Quality of Result, to avoid wasting the computational efforts on noncritical operations. Furthermore, the proposed framework supports global code motion, which helps to improve the speed performance of the hardware implementation by distributing the execution time of operations across the their parent BB. Hongbin Zheng, Qingrui Liu, Dihu Chen |
ASP-DAC | 4 |
| 2013 | hJam: Attachment Transmission in WLANsabstractEffective coordination can dramatically reduce radio interference and avoid packet collisions for multistation wireless local area networks (WLANs). Coordination itself needs consume communication resource and thus competes with data transmission for the limited wireless radio resources. In traditional approaches, control frames and data packets are transmitted in an alternate manner, which brings a great deal of coordination overhead. In this paper, we propose a new communication model where the control frames can be "attachedâ to the data transmission. Thus, control messages and data traffic can be transmitted simultaneously and consequently the channel utilization can be improved significantly. We implement the idea in OFDM-based WLANs called hJam, which fully explores the physical layer features of the OFDM modulation method and allows one data packet and a number of control messages to be transmitted together. hJam is implemented on the GNU Radio testbed consisting of eight USRP2 nodes. We also conduct comprehensive simulations and the experimental results show that hJam can improve the WLANs efficiency by up to 200 percent compared with the existing 802.11 family protocols. Kaishun Wu, Haochao Li, Lu Wang 0002, Youwen Yi, Yunhuai Liu, Dihu Chen, Qian Zhang 0001, Lionel M. Ni |
IEEE Trans. Mob. Comput. | 6 |
| 2013 | CSI-Based Indoor LocalizationabstractIndoor positioning systems have received increasing attention for supporting location-based services in indoor environments. WiFi-based indoor localization has been attractive due to its open access and low cost properties. However, the distance estimation based on received signal strength indicator (RSSI) is easily affected by the temporal and spatial variance due to the multipath effect, which contributes to most of the estimation errors in current systems. In this work, we analyze this effect across the physical layer and account for the undesirable RSSI readings being reported. We explore the frequency diversity of the subcarriers in orthogonal frequency division multiplexing systems and propose a novel approach called FILA, which leverages the channel state information (CSI) to build a propagation model and a fingerprinting system at the receiver. We implement the FILA system on commercial 802.11 NICs, and then evaluate its performance in different typical indoor scenarios. The experimental results show that the accuracy and latency of distance calculation can be significantly enhanced by using CSI. Moreover, FILA can significantly improve the localization accuracy compared with the corresponding RSSI approach. Kaishun Wu, Jiang Xiao 0001, Youwen Yi, Dihu Chen, Lionel M. Ni |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2011 | 1024-point pipeline FFT processor with pointer FIFOs based on FPGAabstractDesign and optimized implementation of a 16-bit and 32-bit 1024-point pipeline FFT processor is presented in this paper. The architecture of the FFT is based on R22SDF algorithm with new pointer FIFO embedded with gray code counters. It is implemented in Spartan-3E, Spartan-6 and Virtex-4 devices and fully tested by method of co-simulation using SMIMS®VeriLink®as a bridge that connects software(Matlab®Simulink®) and real hardware-FPGA targets. The implementation results show that our pointer FIFO FFT processor could use lower resource, but achieve higher performance. Our 16-bit 1024-point FFT processor only costs 2580 slices, 2030 slice flip flops and just 2 block RAMs, achieving the maximum clock frequency of 92.6 MHz with the throughput per area of 0.035 Msamples/s/area. Due to the parameterized input wordlength, output wordlength, Twiddle Factors wordlength and processing stages, it is easily to implement a 16-point, 64-point, 256-point, 1024-point,4096-point or higher power of 4 points pointer FIFO FFT processor synthesized from the same code just through modifying the corresponding parameters. Guanwen Zhong, Hongbin Zheng, ZhenHua Jin, Dihu Chen, Zhiyong Pang |
VLSI-SoC | 4 |