EDBT 2026 Demo / reviewers in the wild / expert
Ronghao Lin
dblp:255/5811
· DBLP profile ↗
23ranked-venue papers
10as first author
22since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Security and privacy · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Joint CFO estimation and interference suppression for Multi-Static sensing in ISAC systems
Zhuoli Liu, Ronghao Lin, Hing-Cheung So, Jian Li 0001 |
Signal Process. | 2 |
| 2026 | Affection-Guided Bottleneck Diffusion for Missing Modality Issue in Multimodal Affective ComputingabstractMissing modality issue in multimodal affective computing severely hinders the robustness and performance of multimodal learning, particularly in real-world scenarios. Existing methods often fail in fully exploiting the remaining modalities, leading to noisy reconstruction process for the missing modalities and yielding suboptimal results. Besides, most of these methods rely on designing sophisticated networks to handle various missing scenarios, which prevents them from taking advantage of the original pre-trained multimodal networks trained for complete multimodal inputs. To address these challenges, we propose Affection-guided Bottleneck Diffusion (ABDiff), a novel approach leveraging score-based diffusion generative encoders to reconstruct missing modalities in the latent space without modification to the pre-trained fusion models. By incorporating self- and cross-attention mechanisms inside and among the missing and remaining modalities, ABDiff captures both modality-specific dynamics and cross-modal interactions during generation. Furthermore, an affection-guided information bottleneck is introduced to filter task-unrelated noise and modality-specific redundancy, stabilizing the generation process of missing modalities. The generated representations are seamlessly integrated with the remaining modalities into the pre-trained fusion networks. Extensive experiments on four public multimodal affective computing datasets demonstrate that ABDiff surpasses previous methods under both complete and incomplete modality scenarios. The code is released inhttps://github.com/RH-Lin/ABDiff. Ronghao Lin, Qiaolin He, Sijie Mai, Haifeng Hu 0001 |
IEEE Trans. Affect. Comput. | 1 |
| 2025 | Diff-LMM: Diffusion Teacher-Guided Spatio-Temporal Perception for Video Large Multimodal ModelsabstractDynamic spatio-temporal understanding is essential for video-based multimodal tasks, yet existing methods often struggle to capture fine-grained temporal and spatial relationships in long videos. Current approaches primarily rely on pre-trained CLIP encoders, which excel in semantic understanding but lack spatially-aware visual context. This leads to hallucinated results when interpreting fine-grained objects or scenes. To address these limitations, we propose a novel framework that integrates diffusion models into multimodal video models. By employing diffusion encoders at intermediate layers, we enhance visual representations through feature alignment and knowledge distillation losses, significantly improving the model's ability to capture spatial patterns over time. Additionally, we introduce a multi-level alignment strategy to learn robust feature correspondence from pre-trained diffusion models. Extensive experiments on benchmark datasets demonstrate our approach's state-of-the-art performance across multiple video understanding tasks. These results establish diffusion models as a powerful tool for enhancing multimodal video models in complex, dynamic scenarios. Jisheng Dang, Ligen Chen, Jingze Wu, Ronghao Lin, Bimei Wang, Yun Wang 0053, Nannan Zhu, Teng Wang 0007 |
IJCAI | 4 |
| 2025 | E3RG: Building Explicit Emotion-driven Empathetic Response Generation System with Multimodal Large Language Model
Ronghao Lin, Shuai Shen, Weipeng Hu, Qiaolin He, Aolin Xiong, Haifeng Hu 0001, Yap-Peng Tan |
ACM Multimedia | 1 |
| 2025 | CyIN: Cyclic Informative Latent Space for Bridging Complete and Incomplete Multimodal LearningabstractMultimodal machine learning, mimicking the human brain’s ability to integrate various modalities has seen rapid growth. Most previous multimodal models are trained on perfectly paired multimodal input to reach optimal performance. In real‑world deployments, however, the presence of modality is highly variable and unpredictable, causing the pre-trained models in suffering significant performance drops and fail to remain robust with dynamic missing modalities circumstances. In this paper, we present a novel Cyclic INformative Learning framework (CyIN) to bridge the gap between complete and incomplete multimodal learning. Specifically, we firstly build an informative latent space by adopting token- and label-level Information Bottleneck (IB) cyclically among various modalities. Capturing task-related features with variational approximation, the informative bottleneck latents are purified for more efficient cross-modal interaction and multimodal fusion. Moreover, to supplement the missing information caused by incomplete multimodal input, we propose cross-modal cyclic translation by reconstruct the missing modalities with the remained ones through forward and reverse propagation process. With the help of the extracted and reconstructed informative latents, CyIN succeeds in jointly optimizing complete and incomplete multimodal learning in one unified model. Extensive experiments on 4 multimodal datasets demonstrate the superior performance of our method in both complete and diverse incomplete scenarios. Ronghao Lin, Qiaolin He, Sijie Mai, Aolin Xiong, Yap-Peng Tan, Haifeng Hu 0001 |
NeurIPS | 1 |
| 2025 | Prompt-Guided Transformer and MLLM Interactive Learning for Text-Based Pedestrian SearchabstractAiming to retrieve pedestrian images based on a textual description query, Text-Based Pedestrian Search (TBPS) gains increasingly attention due to its applications in security surveillance. As a fine-grained classification task, TBPS requires identifying images of individuals with different semantic contexts yet the same identity, as well as distinguishing images of individuals who share similar appearances but distinct identities. Consequently, TBPS is challenged by semantic variations in positive pairs and appearance similarity between negative pairs. To tackle these challenges, we propose the Prompt-guided Transformer and MLLM Interactive learning (PTMI) model to learn identity-discriminative representations across different modalities. PTMI consists of three components: the Prompt-guided Transformer (Promformer), MLLM Interactive Learning (MIL) and Dual-branch Cross-modal Learning (DCL). Firstly, the Promformer is designed to handle semantic variations in positive pairs by introducing learnable prompts, composing of three types: instance-shared, instance-specific and layer-specific. Optimized by Cross-modal Intra-class Consistency (CIC) loss, these prompts minimize intra-class variations and retrieve positive images with various semantics. Secondly, the MIL component is introduced to address appearance similarity between negative pairs by focusing on key image patches and description words filtering by the local discriminator. Powered by Multimodal Large Language Model (MLLM), the local discriminator adopts soft attention to highlight important image regions and descriptive words, which preserves semantic information while emphasize discriminative details. Lastly, the DCL integrates global and local branches to bridge modality discrepancies. The global branch employs SDM loss for heterogeneous distribution alignment, while the local branch applies Anchor-Based Contrastive (ABC) loss for instance-level contrastive learning. Unlike conventional contrastive loss, ABC loss leverages MLLM features as anchors to decouple modality and semantic differences, enhancing alignment efficiency. Extensive experiments on three TBPS datasets have validated the effectiveness of PTMI. Zefeng Lu, Ronghao Lin, Yap-Peng Tan, Haifeng Hu 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Disentangling Modality and Posture Factors: Memory-Attention and Orthogonal Decomposition for Visible-Infrared Person Re-IdentificationabstractStriving to match the person identities between visible (VIS) and near-infrared (NIR) images, VIS-NIR reidentification (Re-ID) has attracted increasing attention due to its wide applications in low-light scenes. However, owing to the modality and pose discrepancies exhibited in heterogeneous images, the extracted representations inevitably comprise various modality and posture factors, impacting the matching of cross-modality person identity. To solve the problem, we propose a disentangling modality and posture factors (DMPFs) model to disentangle modality and posture factors by fusing the information of features memory and pedestrian skeleton. Specifically, the DMPF comprises three modules: three-stream features extraction network (TFENet), modality factor disentanglement (MFD), and posture factor disentanglement (PFD). First, aiming to provide memory and skeleton information for modality and posture factors disentanglement, the TFENet is designed as a three-stream network to extract VIS-NIR image features and skeleton features. Second, to eliminate modality discrepancy across different batches, we maintain memory queues of previous batch features through the momentum updating mechanism and propose MFD to integrate features in the whole training set by memory-attention layers. These layers explore intramodality and intermodality relationships between features from the current batch and memory queues under the optimization of the optimal transport (OT) method, which encourages the heterogeneous features with the same identity to present higher similarity. Third, to decouple the posture factors from representations, we introduce the PFD module to learn posture-unrelated features with the assistance of the skeleton features. Besides, we perform subspace orthogonal decomposition on both image and skeleton features to separate the posture-related and identity-related information. The posture-related features are adopted to disentangle the posture factors from representations by a designed posture-features consistency (PfC) loss, while the identity-related features are concatenated to obtain more discriminative identity representations. The effectiveness of DMPF is validated through comprehensive experiments on two VIS-NIR pedestrian Re-ID datasets. Zefeng Lu, Ronghao Lin, Haifeng Hu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | On Binary Sequence Design via PSL MinimizationabstractWe introduce an efficient gradient based algorithm to minimize the peak sidelobe level (PSL) for binary sequence designs with or without low correlation zone requirements. The proposed algorithm calculates the effective step sizes by leveraging fast Fourier transform operations and employing low-complexity updates, and its local convergence is proved. Numerical examples are provided to demonstrate that our approach can outperform the state-of-the-art algorithms in generating binary sequences with lower PSL values. Ronghao Lin, Hing-Cheung So, Jian Li 0001 |
IEEE Signal Process. Lett. | 1 |
| 2024 | Multi-Task Momentum Distillation for Multimodal Sentiment AnalysisabstractIn the field of Multimodal Sentiment Analysis (MSA), the prevailing methods are devoted to developing intricate network architectures to capture the intra- and inter-modal dynamics, which necessitates numerous parameters and poses more difficulties in terms of interpretability in multimodal modeling. Besides, the heterogeneous nature of multiple modalities (text, audio, and vision) introduces significant modality gaps, thereby making multimodal representation learning an ongoing challenge. To address the aforementioned issues, by considering the learning process of modalities as multiple subtasks, we propose a novel approach named Multi-Task Momentum Distillation (MTMD) which succeeds in reducing the gap among different modalities. Specifically, according to the abundance of semantic information, we treat the subtasks of textual and multimodal representations as the teacher networks while the subtasks of acoustic and visual representations as the student ones to present knowledge distillation, which transfers the sentiment-related knowledge guided by the regression and classification subtasks. Additionally, we adopt unimodal momentum models to explore modality-specific knowledge deeply and employ adaptive momentum fusion factors to learn a robust multimodal representation. Furthermore, we provide a theoretical perspective of mutual information maximization by interpreting MTMD as generating sentiment-related views in various ways. Extensive experiments illustrate the superiority of our approach compared with the state-of-the-art methods in MSA. Ronghao Lin, Haifeng Hu 0001 |
IEEE Trans. Affect. Comput. | 1 |
| 2024 | Mind the Inconsistent Semantics in Positive Pairs: Semantic Aligning and Multimodal Contrastive Learning for Text-Based Pedestrian SearchabstractAiming at retrieving pedestrian images based on a provided textual description query, Text-Based Pedestrian Search (TBPS) has gained attention due to its implications in public security tasks such as suspect tracking. Nevertheless, the modality discrepancies between textual descriptions and visual images pose a challenge in aligning semantic information between these two modalities. Moreover, the text description annotated on a particular pedestrian image may not align with the content of other images sharing the same identity, due to variations in viewpoint. These text-image pairs exhibiting inconsistent semantics, termed weak positive pairs, have a discernible impact on the model’s performance. To address these challenges, we propose a Semantic Aligning and Multimodal Contrastive learning (SAMC) model to capture cross-modality identity-invariant features, including three modules: Multi-modality Features Fusion (MFF), Semantic-aligning Optimal Transport (SOT), and Multi-modality Contrastive Learning (MCL). Firstly, the MFF is designed to fuse textual and visual information and extract identity-discriminative multimodal features using self- and cross-attention mechanisms. The multimodal features act as anchors, bridging the gap between the two modalities and enhancing the identity-invariance of unimodal features. Secondly, the SOT is designed to address the semantic misalignment issue between textual descriptions and visual images. Utilizing the Optimal Transport (OT) theory, SOT encourages high features similarity between positive samples from different modalities, thereby exploring semantic relationships between image and text data without requiring extra supervised labels. Lastly, the MCL is introduced to narrow the modality gap, compelling two unimodal features towards the identity-discriminative multimodal features through contrastive learning. Different temperature coefficients are employed for strong and weak positive pairs to mitigate the inconsistency in text-image pair correlation. The effectiveness of SAMC is validated by extensive comprehensive experiments on three TBPS datasets. Zefeng Lu, Ronghao Lin, Haifeng Hu 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | A Robust Database Watermarking Scheme That Preserves Statistical CharacteristicsabstractDatabase watermarking can be used for copyright verification and leakage traceability, effectively protecting the security of the database. However, the existing watermarking schemes commonly embed watermarks by modifying the original data, which changes the statistical characteristics and affects the statistical analysis of the database. Therefore, this paper proposes SCPW, aStatisticalCharacteristicsPreserving robust databaseWatermarking framework. First, we perform a theoretical analysis and propose a data modification scheme maintaining the statistical characteristics unchanged. Then, we establish the correspondence between the data and the watermarks that need to be embedded in it by grouping. Finally, the watermark message is embedded into the database through data verification and modification. Specifically, for data that needs to be watermarked, we first verify whether the potential watermark bits extracted from the data are the same as bits that need to be embedded. If they are the same, we regard this original data, usually a floating point number, as a “good number” and do not modify it. Otherwise, we modify the data until it becomes a “good number” using a data modification scheme that preserves the statistical characteristics proposed by the theoretical analysis. In addition, we also use the genetic algorithm to optimize the grouping results and increase the proportion of “good number”, thereby reducing the proportion of data that needs to be modified and further reducing distortion. To our best knowledge, SCPW is the first watermarking scheme that ensures the preservation of statistical characteristics, and the experimental results also prove its excellent ability to preserve statistical characteristics compared to existing schemes. Moreover, experiments also illustrate that our method is robust against a wide range of attacks. When under deletion attack (deletion rate = 90%), the bit error rate of watermark extraction is only 0.8%, which is more than 12% lower than the current best method. Zhiwen Ren, Han Fang 0004, Jie Zhang 0073, Zehua Ma, Ronghao Lin, Weiming Zhang 0001, Nenghai Yu |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2024 | Dynamically Shifting Multimodal Representations via Hybrid-Modal Attention for Multimodal Sentiment AnalysisabstractIn the field of multimodal machine learning, multimodal sentiment analysis task has been an active area of research. The predominant approaches focus on learning efficient multimodal representations containing intra- and inter-modality information. However, the heterogeneous nature of different modalities brings great challenges to multimodal representation learning. In this article, we propose a multi-stage fusion framework to dynamically fine-tune multimodal representations via a hybrid-modal attention mechanism. Previous methods mostly only fine-tune the textual representation due to the success of large corpus pre-trained models and neglect the inconsistency problem of different modality spaces. Thus, we design a module called the Multimodal Shifting Gate (MSG) to fine-tune the three modalities by modeling inter-modality dynamics and shifting representations. We also adopt a module named Masked Bimodal Adjustment (MBA) on the textual modality to improve the inconsistency of parameter spaces and reduce the modality gap. In addition, we utilize syntactic-level and semantic-level textual features output from different layers of the Transformer model to sufficiently capture the intra-modality dynamics. Moreover, we construct a Shifting HuberLoss to robustly introduce the variation of the shifting value into the training process. Extensive experiments on the public datasets, including CMU-MOSI and CMU-MOSEI, demonstrate the efficacy of our approach. Ronghao Lin, Haifeng Hu 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | Tri-Level Modality-Information Disentanglement for Visible-Infrared Person Re-IdentificationabstractAiming to match the person identity between daytime VISible (VIS) and nighttime Near-InfraRed (NIR) images, VIS-NIR re-identification (Re-ID) has attracted increasing attention due to its wide applications in low-light scenes. However, dramatic modality discrepancies between VIS and NIR images lead to a considerable intra-class gap in the feature space, which impacts identity matching. To bridge the modality gap, we propose a Tri-level Modality-information Disentanglement (TMD) to disentangle modality information at the levels of raw image, features distribution and instance features. Our model consists of three key modules, including Style-Aligned Converter (SAC), Two-Steps Wasserstein Loss (TSWL) and Self-supervised Orthogonal Disentanglement (SOD) to handle the modality information at the three levels. Firstly, aiming at reducing modality discrepancy at image-level, the SAC is introduced to generate style-aligned images by the designed style converter and$\mathcal {A}$-distance learning approach. The SAC can effectively alleviate the style discrepancy between VIS and NIR images with a negligible increase in model complexity. Secondly, considering the heterogeneity of VIS and NIR feature distribution caused by the structure- and style-misaligned raw images, we propose the TSWL to decrease the VIS-NIR gap at distribution-level by two distribution alignment steps. Specifically, after generating style-consistent images, we eliminate modality-related discrepancy by aligning the distribution between structure-aligned original and generated VIS/NIR images and bridge the modality-unrelated gap by aligning the style-consistent generated VIS-NIR images. Thirdly, focusing on further reducing the modality discrepancy at instance-level, the SOD is presented to construct orthogonal constraints between the extracted modality- and identity-related features. Since the modality-related factors are disentangled from the instance features, the proposed TMD efficiently learns the modality-unrelated and identity-discriminative representations, which are productive to conduct person Re-ID task on the VIS-NIR images. Comprehensive experiments are carried out on two cross-modality pedestrian Re-ID datasets to demonstrate the effectiveness of TMD. Zefeng Lu, Ronghao Lin, Haifeng Hu 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Efficient binary sequence set designs for MIMO PMCW radar
Yuanbo Cheng, Ronghao Lin, Jian Li 0001, Hing-Cheung So |
Signal Process. | 3 |
| 2023 | Effective and efficient gradient based methods for low-bit discrete-phase sequence designs
Ronghao Lin, Xiaolei Shang, Jian Li 0001 |
Signal Process. | 1 |
| 2023 | MissModal: Increasing Robustness to Missing Modality in Multimodal Sentiment AnalysisabstractAbstract When applying multimodal machine learning in downstream inference, both joint and coordinated multimodal representations rely on the complete presence of modalities as in training. However, modal-incomplete data, where certain modalities are missing, greatly reduces performance in Multimodal Sentiment Analysis (MSA) due to varying input forms and semantic information deficiencies. This limits the applicability of the predominant MSA methods in the real world, where the completeness of multimodal data is uncertain and variable. The generation-based methods attempt to generate the missing modality, yet they require complex hierarchical architecture with huge computational costs and struggle with the representation gaps across different modalities. Diversely, we propose a novel representation learning approach named MissModal, devoting to increasing robustness to missing modality in a classification approach. Specifically, we adopt constraints with geometric contrastive loss, distribution distance loss, and sentiment semantic loss to align the representations of modal-missing and modal-complete data, without impacting the sentiment inference for the complete modalities. Furthermore, we do not demand any changes in the multimodal fusion stage, highlighting the generality of our method in other multimodal learning systems. Extensive experiments demonstrate that the proposed method achieves superior performance with minimal computational costs in various missing modalities scenarios (flexibility), including severely missing modality (efficiency) on two public MSA datasets. Ronghao Lin, Haifeng Hu 0001 |
Trans. Assoc. Comput. Linguistics | 1 |
| 2023 | Modality and Camera Factors Bi-Disentanglement for NIR-VIS Object Re-IdentificationabstractAiming to match object identities across different modalities of images, the challenging task namely cross-modality object re-identification (NIR-VIS object Re-ID), has attracted increasing attention due to its wide application in low-light scenes. However, dramatic modality-dependent and camera-related discrepancies between Near-InfraRed-spectrum (NIR) and VISible-spectrum (VIS) images lead to a considerable intra-class gap in feature space. To address the problem, we propose a novel Modality and Camera factors Bi-Disentanglement (MCBD) model to learn modality-independent and camera-unrelated features for NIR-VIS object Re-ID. Our model consists of three key modules, including Confused Modality Generation (CMG), Modality-independent Information Distillation (MID), and Cameras Factor Disentanglement (CFD). Firstly, aiming at aligning image style between NIR and VIS data, the CMG utilizes a designed channel-interactive generator to generate confused modality images which preserves the structure information of original images. Besides, CMG is trained with confused adversarial learning which bridges the modality gap at the image level. Nevertheless, training the model with the confused modality images discards identity-related information such as color and contrast, which is not conducive to the extraction of distinctive features. To solve this problem, the MID is presented to distill out modality-independent information by feeding the original images to the model and reconstructing corresponding NIR and VIS modalities features to the confused modality images. Finally, due to the complexity of camera-related information in images, identity representation inevitably contains camera-related elements such as background and perspective information, which may interfere the matching process. To address this issue, the CFD is introduced to disentangle camera-related factors by the designed three-streams network and two factor decoupling losses, i.e., Camera-Camera Factor loss (CCF) and Identity-Camera Factor (ICF) loss. Comprehensive experiments are carried out on two cross-modality pedestrian Re-ID datasets and a cross-modality vehicle Re-ID dataset to demonstrate that the MCBD is effective in cross-modality object Re-ID task. Zefeng Lu, Ronghao Lin, Haifeng Hu 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2023 | MART: Mask-Aware Reasoning Transformer for Vehicle Re-IdentificationabstractAs a significant topic in Intelligent Transportation Systems (ITS), vehicle Re-Identification (Re-ID) has attracted increasing research attention. However, the variation of shooting scenes and the similar appearance among the vehicles with the same type and color lead to large intra-class variances and small inter-class variances, respectively. To address the problems, we propose a novel Mask-Aware Reasoning Transformer (MART) to extract the background-unrelated global features and perspective-invariant local features. The MART contains three effective modules including Foreground Global Features Extraction (FGFE), Mask-guided Local Features Extraction (MLFE) and Cross-images Local Features Reasoning (CLFR). Firstly, due to the complexity of background information in images, identity representation inevitably contains background elements, which may impact identity matching. To address this issue, we propose the FGFE to extract background-independent global features by introducing the mask semantic information to the inputs of Vision Transformer (ViT). Secondly, to fill the gap that the previous local features extraction methods cannot be directly applied to ViT, the MLFE is presented to extract distinctive local features by recombining token features according to vehicle mask. Thirdly, when local components are invisible in the image due to the occlusion problem, the corresponding local information is absent from the image, leading to unreliable local features. To solve this problem, the CLFR is proposed to reason the occluded local features by exploiting the correlation between cross-image local features. We carry out comprehensive experiments to illustrate the effectiveness of the MART on two challenging datasets. Zefeng Lu, Ronghao Lin, Haifeng Hu 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Mask-Aware Pseudo Label Denoising for Unsupervised Vehicle Re-IdentificationabstractAs a significant part of Intelligent Transportation System (ITS), vehicle Re-Identification (Re-ID) aims to retrieve all target vehicle images captured from non-overlapping cameras. Though the Re-ID methods based on supervised learning have achieved rapid progress, they are still difficult to be applied in real scenarios due to the domain bias between the training set and real scenarios. Recently, methods based on unsupervised learning have been proposed to address the problem of domain bias by exploring techniques of pseudo-label generation. However, these methods suffer from pseudo-label noise. To solve this problem, we propose the Mask-Aware Pseudo Label Denoising framework (MAPLD) consisting of three key components, i.e., Mask-Aware Feature Extraction (MAFE), Adaptive Threshold Neighborhood Consistency (ATNC), and Compact Loss (CL). Firstly, the MAFE is proposed to improve the distinguishability of feature representation and widen the gap in feature space among vehicles with different IDs. Next, the ATNC is introduced to filter out pseudo-label noise of hard negative samples by comparing the image ID of the samples in their neighborhood set i.e., neighborhood consistency. Moreover, the threshold of neighborhood consistency is adaptively adjusted according to feature similarity ranking, which is robust to hyper-parameter variation. Finally, consisting of regression term and compact term, the CL is designed to drive the cluster more compact and alleviate the impact of outliers of hard positive samples. Extensive experiments on VeRi-776 and VeRi-Wild datasets demonstrate that MAPLD can generate reliable pseudo-labels and achieve superior performance in unsupervised target-only and unsupervised domain adaptation tasks. Zefeng Lu, Ronghao Lin, Qiaolin He, Haifeng Hu 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | On the design of linear sparse arrays with beampattern shift invariant properties
Ronghao Lin, Gang Wang 0032, Jian Li 0001 |
Signal Process. | 1 |
| 2022 | Identity-Unrelated Information Decoupling Model for Vehicle Re-IdentificationabstractAs an indispensable part of intelligent transportation system (ITS), vehicle re-identification (Re-ID) aims to retrieve all target vehicle images captured from non-overlapping cameras. However, this task remains very challenging due to the variation of camera perspective and the similar appearance among the vehicles with the same type and color, i.e., large intra-class variances and small inter-class variances. Previous methods have made a great progress on vehicle Re-ID by leveraging local details and aligning local features, this issue is still far from being solved. In this work, we propose to decouple identity-unrelated information from vehicle representation, tackling the problems of camera perspective variation and vehicle appearance similarity. The keypoint of this method is to learn a distinguishable feature embedding that is independent of identity-unrelated information. Specially, the novel Identity-Unrelated Information Decoupling (IUID) paradigm is designed to learn invariant features of the vehicle with the same ID in different scenes. In our approach, identity-unrelated information can be divided into two kinds of information, i.e., camera perspective information and background information. For the former, through a feature-level camera generative adversarial module, we can decouple camera perspective information from the feature embedding after extracting invariant features across different cameras perspective. For the latter, we propose a vehicle-mask transformer to enhance the attention of the model on local details while reducing the impact of background information. Extensive experiments on two public datasets demonstrate the superiority of IUID over the current state-of-the-arts methods. Zefeng Lu, Ronghao Lin, Xulei Lou, Lifeng Zheng, Haifeng Hu 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Efficient Long Periodic Binary Sequence Designs for Automotive RadarabstractPeriodic binary sequences with good autocorrelation properties are useful for phase-modulated continuous wave (PMCW) automotive radar systems. The optimization problem encountered in the periodic binary sequence designs is NP-hard, and most of the existing algorithms are not suitable for designing long periodic sequences due to their heavy computational burdens. We consider an FFT-based algorithm for the efficient designs of long periodic binary sequences with arbitrary period lengths and ample diversity. Several numerical examples are provided to demonstrate the performance of the proposed algorithm. Ronghao Lin, Jian Li 0001 |
ICASSP | 2 |
| 2020 | On Binary Sequence Set Design with Applications to Automotive RadarabstractWe consider herein the case of two vehicles equipped with multi-input multi-output (MIMO) automotive radars driving next to each other. We assume that 5G communications allow us to coordinate the radar probing waveforms for the vehicles. Then the binary sequence sets transmitted by different vehicles should meet the requirement that the cross-correlations of the sequence sets between vehicles are as low as possible for all time lags, while the auto-correlation sidelobes and cross-correlations within a low correlation zone (LCZ) of the binary sequence sets transmitted by each vehicle are lower than a desired level. We establish an optimization problem to realize these goals. We consider the coordinate-descent (CD) framework and we solve the optimization problem efficiently by taking full advantage of the fast Fourier transforms (FFTs) and introducing computationally efficient updating procedures within the CD iterations. Numerical examples are provided to demonstrate that the proposed algorithms can be used to effectively and efficiently design binary sequence sets, including long sequence sets, useful for MIMO automotive radar applications. Ronghao Lin, Jian Li 0001 |
ICASSP | 1 |