VLDB 2026 Research / reviewers in the wild / expert
Heming Du
dblp:244/8133
· DBLP profile ↗
30ranked-venue papers
4as first author
27since 2021 · last 2026
0000-0002-7391-0449ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 3 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FreDN: Spectral Disentanglement for Time Series Forecasting via Learnable Frequency DecompositionabstractTime series forecasting is essential in a wide range of real world applications. Recently, frequency-domain methods have attracted increasing interest for their ability to capture global dependencies. However, when applied to non-stationary time series, these methods encounter the spectral entanglement and the computational burden of complex-valued learning. The spectral entanglement refers to the overlap of trends, periodicities, and noise across the spectrum due to spectral leakage and the presence of non-stationarity. However, existing decompositions are not suited to resolving spectral entanglement. To address this, we propose the Frequency Decomposition Network (FreDN), which introduces a learnable Frequency Disentangler module to separate trend and periodic components directly in the frequency domain. Furthermore, we propose a theoretically supported ReIm Block to reduce the complexity of complex-valued operations while maintaining performance. We also re-examine the frequency-domain loss function and provide new theoretical insights into its effectiveness. Extensive experiments on seven long-term forecasting benchmarks demonstrate that FreDN outperforms state-of-the-art methods by up to 10%. Furthermore, compared with standard complex-valued architectures, our real-imaginary shared-parameter design reduces the parameter count and computational cost by at least 50%. Zhongde An, Jinhong You, Jiyanglin Li, Heming Du, Shouguo Du |
AAAI | 6 |
| 2026 | Mobile Auslan: A multimodal dialogue-centered sign language learning systemabstractLearning sign language is not only a gateway to linguistic competence but also to cultural participation, self-expression, and meaningful interaction within the Deaf community. Recent advances in sign language education technologies have made notable progress in supporting vocabulary acquisition and sentence-level translation. However, dialogue, as one critical component of natural language use, remains largely absent from existing systems. We identify two underlying challenges that contribute to this gap: the lack of recognition robustness to viewpoint variation, which constrains expressive freedom during signing; and the limited semantic modeling capabilities needed to support discourse-level interpretation and interaction. To address these issues, we propose a learning-centered system that integrates pose-based multi-view augmentation and a multi-agent language modeling workflow. This system supports free-form input across diverse signing perspectives and provides structured feedback across word, sentence, and dialogue levels. Built as a modular and deployable platform, the system demonstrates strong recognition performance and learner engagement across varied input conditions. Through this integration of spatial robustness and semantic scaffolding, our work advances the design of sign language learning technologies toward more interactive, expressive, and pedagogically grounded experiences. However, the current system remains limited by its focus on successive (non-continuous) signing and by its moderate vocabulary and restricted pedagogical scope. Future work will therefore extend lexical coverage, incorporate more natural continuous signing, and investigate richer educational functions. • A View-invariant Pose-based Isolated Auslan Sign Recognition. • A Multi-agent Workflow Providing Transparent Dialogue Support. • An Interactive, Learner-paced Design for Auslan Learning System. Hongwei Sheng, Heming Du, Xin Yu 0002 |
Comput. Vis. Image Underst. | 3 |
| 2026 | Diverse Sign Language TranslationabstractAbstract Like spoken languages, a single sign language expression could correspond to multiple valid textual interpretations. Hence, learning a rigid one-to-one mapping for sign language translation (SLT) models might be inadequate, particularly in the case of limited data. In this work, we introduce a Diverse Sign Language Translation (DivSLT) task, aiming to generate diverse yet accurate translations for sign language videos. Firstly, we employ large language models (LLM) to generate multiple references for the widely-used CSL-Daily and PHOENIX14T SLT datasets. Here, native speakers are only invited to touch up inaccurate references, thus significantly improving the annotation efficiency. Secondly, we provide a benchmark model to spur research in this task. Specifically, we investigate multi-reference training strategies enabling our DivSLT model to achieve diverse translations. Then, to enhance translation accuracy, we employ the max-reward-driven reinforcement learning objective that maximizes the reward of the translated result. Additionally, we utilize multiple metrics to assess the accuracy, diversity, and semantic precision of the DivSLT task. Experimental results on the enriched datasets demonstrate that our DivSLT method achieves not only better translation performance but also diverse translation results. Shaozu Yuan, Heming Du, Xin Yu 0002 |
Int. J. Comput. Vis. | 4 |
| 2025 | M3GYM: A Large-Scale Multimodal Multi-view Multi-person Pose Dataset for Fitness Activity Understanding in Real-world SettingsabstractHuman pose estimation is a critical task in computer vision for applications in sports analysis, healthcare monitoring, and human-computer interaction. However, existing human pose datasets are collected either from custom-configured laboratories with complex devices or they only include data on single individuals, and both types typically capture daily activities. In this paper, we introduce the M3GYM dataset, a large-scale multimodal, multi-view, and multi-person pose dataset collected from a real gym to address the limitations of existing datasets. Specifically, we collect videos for 82 sessions from the gym, each session lasting between 40 to 60 minutes. These videos are gathered by 8 cameras, including over 50 subjects and 47 million frames. These sessions include 51 Normal fitness exercise sessions as well as 17 Pilates and 14 Yoga sessions. The exercises cover a wide range of poses and typical fitness activities, particularly in Yoga and Pilates, featuring poses with stretches, bends, and twists, e.g., humble warrior, fire hydrants and knee hover side twists. Each session involves multiple subjects, leading to significant self-occlusion and mutual occlusion in single views. Moreover, the gym has two symmetric floor mirrors, a feature not seen in previous datasets, and seven lighting conditions. We provide frame-level multimodal annotations, including 2D&3D keypoints, subject IDs, and meshes. Additionally, M3GYM uniquely offers labels for over 500 actions along with corresponding assessments from sports experts. We benchmark a variety of state-of-the-art methods for several tasks, i.e., 2D human pose estimation, single-view and multi-view 3D human pose estimation, and human mesh recovery. To simulate real-world applications, we also conduct cross-domain experiments across Normal, Yoga, and Pilates sessions. The results show that M3GYM significantly improves model generalization in complex real-world settings. The project is available here. Qingzheng Xu, Ru Cao, Heming Du, Sen Wang 0001, Xin Yu 0002 |
CVPR | 4 |
| 2025 | LDPose: Towards Inclusive Human Pose Estimation for Limb-Deficient Individuals in the Wild
Jiaying Ying, Heming Du, Kaihao Zhang, Lincheng Li, Xin Yu 0002 |
ICCV | 2 |
| 2025 | Quantifying and Narrowing the Unknown: Interactive Text-to-Video Retrieval Via Uncertainty MinimizationabstractDespite recent advances, Text-to-video retrieval (TVR) is still hindered by multiple inherent uncertainties, such as ambiguous textual queries, indistinct text-video mappings, and low-quality video frames. Although interactive systems have emerged to address these challenges by refining user intent through clarifying questions, current methods typically rely on heuristic or ad-hoc strategies without explicitly quantifying these uncertainties, limiting their effectiveness. Motivated by this gap, we propose UMIVR, an Uncertainty-Minimizing Interactive Text-to-Video Retrieval framework that explicitly quantifies three critical uncertainties-text ambiguity, mapping uncertainty, and frame uncertainty-via principled, training-free metrics: semantic entropy-based Text Ambiguity Score (TAS), Jensen-Shannon divergence-based Mapping Uncertainty Score (MUS), and a Temporal Quality-based Frame Sampler (TQFS). By adaptively generating targeted clarifying questions guided by these uncertainty measures, UMIVR iteratively refines user queries, significantly reducing retrieval ambiguity. Extensive experiments on multiple benchmarks validate UMIVR's effectiveness, achieving notable gains in Recall@1 (69.2\% after 10 interactive rounds) on the MSR-VTT-1k dataset, thereby establishing an uncertainty-minimizing foundation for interactive TVR. Bingqing Zhang, Heming Du, Yang Li 0184, Xue Li 0001, Jiajun Liu 0004, Sen Wang 0001 |
ICCV | 3 |
| 2025 | Multimodal Retina Image Analysis Survey: Datasets, Tasks and MethodsabstractRetina images provide a noninvasive view of the central nervous system and microvasculature, making it essential for clinical applications. Changes in the retina often indicate both ophthalmic and systemic diseases, aiding in diagnosis and early intervention. While deep learning algorithms have advanced retina image analysis, a comprehensive review of related datasets, tasks, and benchmarking is still lacking. In this survey, we systematically categorize existing retina image datasets based on their available data modalities, and review the tasks these datasets support in multimodal retina image analysis. We also explain key evaluation metrics used in various retina image analysis benchmarks. By thoroughly examining current datasets and methods, we highlight the challenges and limitations in existing benchmarks and discuss potential research topics in the field. We hope this work will guide future retina analysis methods and promote the shared use of existing data across different tasks. Hongwei Sheng, Heming Du, Sen Wang 0001, Xin Yu 0002 |
IJCAI | 2 |
| 2025 | DAEM: A Decomposed Attention-Enhanced Mamba Model for Multivariate Time Series ForecastingabstractMultivariate time series forecasting (MTSF) is a critical task with applications spanning diverse domains. Effective MTSF requires capturing both inter-series dependencies, which represent the intricate relationships among multiple time series, and intra-series dynamics, which reflect the temporal patterns and variations within individual series. However, the complexity of real-world data, marked by nonlinearity, noise, and dynamic fluctuations, presents significant challenges. In this paper, we propose the Decomposed Attention-Enhanced Mamba Model (DAEM), which combines a learnable decomposition strategy with an attention-based module and the Mamba architecture to effectively capture dynamic trends, seasonal patterns, and the intricate relationships inherent in multivariate time series. By addressing both inter-series dependencies and intra-series dynamics, this innovative design enables DAEM to achieve more accurate and robust time series forecasting. To evaluate the effectiveness of DAEM, we conducted extensive experiments on eight widely used benchmark datasets. The results demonstrate the superiority of our approach, achieving 35 first-place rankings for MSE and 28 first-place rankings for MAE. Heming Du, Jiyanglin Li, Shouguo Du |
IJCNN | 3 |
| 2025 | When One Moment Isn't Enough: Multi-Moment Retrieval with Cross-Moment InteractionsabstractExisting Moment retrieval (MR) methods focus on Single-Moment Retrieval (SMR). However, one query can correspond to multiple relevant moments in real-world applications. This makes the existing datasets and methods insufficient for video temporal grounding.
By revisiting the gap between current MR tasks and real-world applications, we introduce a high-quality datasets called QVHighlights Multi-Moment Dataset (QV-M$^2$), along with new evaluation metrics tailored for multi-moment retrieval (MMR). QV-M$^2$ consists of 2,212 annotations covering 6,384 video segments. Building on existing efforts in MMR, we propose a framework called FlashMMR. Specifically, we propose a Multi-moment Post-verification module to refine the moment boundaries. We introduce constrained temporal adjustment and subsequently leverage a verification module to re-evaluate the candidate segments. Through this sophisticated filtering pipeline, low-confidence proposals are pruned, and robust multi-moment alignment is achieved.
We retrain and evaluate 6 existing MR methods on QV-M$^2$ and QVHighlights under both SMR and MMR settings. Results show that QV-M$^2$ serves as an effective benchmark for training and evaluating MMR models, while FlashMMR provides a strong baseline. Specifically, on QV-M$^2$, it achieves improvements over prior SOTA method by 3.00% on G-mAP, 2.70% on mAP@3+tgt, and 2.56% on mR@3. The proposed benchmark and method establish a foundation for advancing research in more realistic and challenging video temporal grounding scenarios. Code is released at https://github.com/Zhuo-Cao/QV-M2. Heming Du, Bingqing Zhang, Xin Yu 0002, Xue Li 0001, Sen Wang 0001 |
NeurIPS | 2 |
| 2025 | A Multi-Scale Decomposition and Fusion Framework Utilizing Mamba for Enhanced Time Series ForecastingabstractMultivariate time series forecasting presents a significant challenge across various fields, requiring accurate predictions of future values based on multiple interrelated time series. Recent research has shown that the Channel Independent (CI) approach, which processes each sequence independently, can improve prediction accuracy, but neglecting the relation-ships between sequences may result in inadequate generalization. Channel Dependent (CD) methods, while integrating all sequences information, however, may compromise prediction accuracy by mixing potentially unrelated data. In this paper, we propose a novel framework called MDF-Mamba which employs a multi-scale decomposition strategy, utilizing the CI approach at fine scales to capture the unique characteristics of individual sequences, thereby enhancing model robustness. At coarse scales, the CD approach is used to capture correlations between sequences, improving the model’s generalization capabilities. MDF-Mamba fully considers the individual characteristics of sequences and their interrelationships, balancing Channel Independent and Dependent to improve multivariate time series forecasting performance. Extensive experimental results across multiple real-world time series datasets demonstrate that MDF-Mamba achieves state-of-the-art performance. Jiyanglin Li, Heming Du, Shouguo Du, Jinhong You |
SMC | 6 |
| 2025 | FlashVTG: Feature Layering and Adaptive Score Handling Network for Video Temporal GroundingabstractText-guided Video Temporal Grounding (VTG) aims to localize relevant segments in untrimmed videos based on textual descriptions, encompassing two subtasks: Moment Retrieval (MR) and Highlight Detection (HD). Although previous typical methods have achieved commendable results, it is still challenging to retrieve short video moments. This is primarily due to the reliance on sparse and limited decoder queries, which significantly constrain the accuracy of predictions. Furthermore, suboptimal outcomes often arise because previous methods rank predictions based on isolated predictions, neglecting the broader video context. To tackle these issues, we introduce FlashVTG, a framework featuring a Temporal Feature Layering (TFL) module and an Adaptive Score Refinement (ASR) module. The TFL module replaces the traditional decoder structure to capture nuanced video content variations across multiple temporal scales, while the ASR module improves prediction ranking by integrating context from adjacent moments and multi-temporal-scale features. Extensive experiments demonstrate that FlashVTG achieves state-of-the-art performance on four widely adopted datasets in both MR and HD. Specifically, on the QVHighlights dataset, it boosts mAP by 5.8% for MR and 3.3% for HD. For short-moment retrieval, FlashVTG increases mAP to 125% of previous SOTA performance. All these improvements are made without adding training burdens, underscoring its effectiveness. Our code is available at https://github.com/Zhuo-Cao/FlashVTG. Bingqing Zhang, Heming Du, Xin Yu 0002, Xue Li 0001, Sen Wang 0001 |
WACV | 3 |
| 2025 | TokenBinder: Text-Video Retrieval with One-to-Many Alignment ParadigmabstractText-Video Retrieval (TVR) methods typically match query-candidate pairs by aligning text and video features in coarse-grained, fine-grained, or combined (coarse-to-fine) manners. However, these frameworks predominantly employ a one(query)-to-one(candidate) alignment paradigm, which struggles to discern nuanced differences among candidates, leading to frequent mismatches. Inspired by Comparative Judgement in human cognitive science, where decisions are made by directly comparing items rather than evaluating them independently, we propose TokenBinder. This innovative two-stage TVR framework introduces a novel one-to-many coarse-to-fine alignment paradigm, imitating the human cognitive process of identifying specific items within a large collection. Our method employs a Focused-view Fusion Network with a sophisticated cross-attention mechanism, dynamically aligning and comparing features across multiple videos to capture finer nuances and contextual variations. Extensive experiments on six benchmark datasets confirm that TokenBinder substantially outperforms existing state-of-the-art methods. These results demonstrate its robustness and the effectiveness of its fine-grained alignment in bridging intra- and inter-modality information gaps in TVR tasks. Code is avaliable at https://github.com/bingqingzhang/TokenBinder. Bingqing Zhang, Heming Du, Xin Yu 0002, Xue Li 0001, Jiajun Liu 0004, Sen Wang 0001 |
WACV | 3 |
| 2025 | AuslanWeb: A Scalable Web-Based Australian Sign Language Communication System for Deaf and Hearing IndividualsabstractEffective communication between the deaf community and hearing individuals facilitates social inclusion, equal opportunities, and the dignity of vulnerable populations. However, existing region-specific sign language systems are constrained by limited training datasets and narrow topic domains, rendering them ineffective for bridging the linguistic gaps between sign languages and spoken languages. Auslan, as the sign language specific to Australia, still lacks a reliable bidirectional translation tool for effective communication. To address these challenges, we propose AuslanWeb, a web-based system for bidirectional translation of both isolated and successive sign language. For the former, AuslanWeb achieves high-precision mapping between isolated signs (glosses) and spoken language words or phrases through a multimodal recognition system and a versatile Auslan dictionary. For the latter, it leverages the advanced contextual understanding and text generation capabilities of Large Language Models (LLMs) to support bidirectional translation between successive sign language videos and long-form spoken language. By integrating linguistic structure with advanced AI capabilities, AuslanWeb overcomes the limitations of dataset dependency and enhances the scalability of sign language translation systems. The effectiveness of the system is further validated through user feedback, receiving consistent praise from Auslan experts, Australian deaf individuals, and volunteers. The demo video of AuslanWeb is provided here. Heming Du, Hongwei Sheng, Lincheng Li, Kaihao Zhang |
WWW | 2 |
| 2025 | MDAM3: A Misinformation Detection and Analysis Framework for Multitype Multimodal MediaabstractMisinformation is a significant societal issue with potentially severe consequences. It appears in text, image, audio, and video modalities, encompassing various categories such as unimodal deception (fact-conflicting, AI-generated & offensive content) and cross-modal inconsistencies. However, current detection approaches often focus on text and image, overlooking the growing prevalence of misinformation in audio and video content. Moreover, these methods typically tend to address only one or two types of misinformation, failing to address all categories simultaneously. These detectors are also usually designed to make judgments without providing explanations, reducing transparency and limiting their broader applicability. To address these issues, we propose MDAM3, a Misinformation Detection and Analysis Framework for Multitype Multimodal Media. MDAM3 analyzes each input in internal detection and examines relationships across modalities to identify inconsistencies. It utilizes web resources and integrates Large Vision-Language Models (LVLMs) to deliver accurate detection results along with detailed analysis. To evaluate MDAM3, we curate MDAM3-DB, a specialized multitype multimodal misinformation dataset. A user study is conducted to explore MDAM3's usability, interpretability, and effectiveness. We hope this research contributes to advancing misinformation detection methodologies and provides valuable insights for developing robust multimodal analysis tools. Qingzheng Xu, Heming Du, Szymon Lukasik, Tianqing Zhu, Sen Wang 0001, Xin Yu 0002 |
WWW | 2 |
| 2024 | MMOOC: A Multimodal Misinformation Dataset for Out-of-Context News Analysis
Qingzheng Xu, Heming Du, Huiqiang Chen, Bo Liu 0001, Xin Yu 0002 |
ACISP (3) | 2 |
| 2024 | Who is Being Impersonated? Deepfake Audio Detection and Impersonated Identification via Extraction of Id-Specific Features
Tianchen Guo, Heming Du, Huan Huo, Bo Liu 0001, Xin Yu 0002 |
ICA3PP (5) | 2 |
| 2024 | Learning Seasonal-Trend Representations and Conditional Heteroskedasticity for Time Series Analysis
Heming Du, Shouguo Du, Jinhong You |
ICANN (6) | 3 |
| 2024 | MM-WLAuslan: Multi-View Multi-Modal Word-Level Australian Sign Language Recognition DatasetabstractIsolated Sign Language Recognition (ISLR) focuses on identifying individual sign language glosses. Considering the diversity of sign languages across geographical regions, developing region-specific ISLR datasets is crucial for supporting communication and research. Auslan, as a sign language specific to Australia, still lacks a dedicated large-scale word-level dataset for the ISLR task. To fill this gap, we curate \underline{\textbf{the first}} large-scale Multi-view Multi-modal Word-Level Australian Sign Language recognition dataset, dubbed MM-WLAuslan. Compared to other publicly available datasets, MM-WLAuslan exhibits three significant advantages: (1) the largest amount of data, (2) the most extensive vocabulary, and (3) the most diverse of multi-modal camera views. Specifically, we record 282K+ sign videos covering 3,215 commonly used Auslan glosses presented by 73 signers in a studio environment.Moreover, our filming system includes two different types of cameras, i.e., three Kinect-V2 cameras and a RealSense camera. We position cameras hemispherically around the front half of the model and simultaneously record videos using all four cameras. Furthermore, we benchmark results with state-of-the-art methods for various multi-modal ISLR settings on MM-WLAuslan, including multi-view, cross-camera, and cross-view. Experiment results indicate that MM-WLAuslan is a challenging ISLR dataset, and we hope this dataset will contribute to the development of Auslan and the advancement of sign languages worldwide. All datasets and benchmarks are available at MM-WLAuslan. Heming Du, Hongwei Sheng, Hui Chen 0036, Huiqiang Chen, Zhuojie Wu, Xiaobiao Du, Jiaying Ying, Ruihan Lu, Qingzheng Xu, Xin Yu 0002 |
NeurIPS | 2 |
| 2024 | When 3D Bounding-Box Meets SAM: Point Cloud Instance Segmentation with Weak-and-Noisy SupervisionabstractLearning from bounding-boxes annotations has shown great potential in weakly-supervised 3D point cloud instance segmentation. However, we observed that existing methods would suffer severe performance degradation with perturbed bounding box annotations. To tackle this issue, we propose a complementary image prompt-induced weakly-supervised point cloud instance segmentation (CIP-WPIS) method. CIP-WPIS leverages pretrained knowledge embedded in the 2D foundation model SAM and 3D geometric prior to achieve accurate point-wise instance labels from the bounding box annotations. Specifically, CIP-WPIS first selects image views in which 3D candidate points of an instance are fully visible. Then, we generate complementary background and foreground prompts from projections to obtain SAM 2D instance mask predictions. According to these, we assign the confidence values to points indicating the likelihood of points belonging to the instance. Furthermore, we utilize 3D geometric homogeneity provided by superpoints to decide the final instance label assignments. In this fashion, we achieve high-quality 3D point-wise instance labels. Extensive experiments on both Scannet-v2 and S3DIS benchmarks proves that our method not only achieves state-of-the-art performance for bounding-boxes supervised point cloud instance segmentation, but also exhibits robustness against noisy 3D bounding-box annotations. Qingtao Yu, Heming Du, Chen Liu 0028, Xin Yu 0002 |
WACV | 2 |
| 2024 | M3A: A multimodal misinformation dataset for media authenticity analysisabstractWith the development of various generative models, misinformation in news media becomes more deceptive and easier to create, posing a significant problem. However, existing datasets for misinformation study often have limited modalities, constrained sources, and a narrow range of topics. These limitations make it difficult to train models that can effectively combat real-world misinformation. To address this, we propose a comprehensive, large-scale Multimodal Misinformation dataset for Media Authenticity Analysis ( M 3 A ), featuring broad sources and fine-grained annotations for topics and sentiments. To curate M 3 A , we collect genuine news content from 60 renowned news outlets worldwide and generate fake samples using multiple techniques. These include altering named entities in texts, swapping modalities between samples, creating new modalities, and misrepresenting movie content as news. M 3 A contains 708K genuine news samples and over 6M fake news samples, spanning text, images, audio, and video. M 3 A provides detailed multi-class labels, crucial for various misinformation detection tasks, including out-of-context detection and deepfake detection. For each task, we offer extensive benchmarks using state-of-the-art models, aiming to enhance the development of robust misinformation detection systems. • We present M 3 A , a large-scale multimodal misinformation dataset with diverse news samples. • M 3 A includes texts, images, audio, and videos from multiple reputable news outlets. • M 3 A addresses limitations in existing datasets in misinformation generation methods and scale. • We provide multi-class annotations in M 3 A for various key tasks in misinformation detection. • We propose benchmarks for M 3 A using state-of-the-art models and out-of-distribution testing. Qingzheng Xu, Huiqiang Chen, Heming Du, Hu Zhang 0005, Szymon Lukasik, Tianqing Zhu, Xin Yu 0002 |
Comput. Vis. Image Underst. | 3 |
| 2023 | SEFormer: Structure Embedding Transformer for 3D Object DetectionabstractEffectively preserving and encoding structure features from objects in irregular and sparse LiDAR points is a crucial challenge to 3D object detection on the point cloud. Recently, Transformer has demonstrated promising performance on many 2D and even 3D vision tasks. Compared with the fixed and rigid convolution kernels, the self-attention mechanism in Transformer can adaptively exclude the unrelated or noisy points and is thus suitable for preserving the local spatial structure in the irregular LiDAR point cloud. However, Transformer only performs a simple sum on the point features, based on the self-attention mechanism, and all the points share the same transformation for value. A such isotropic operation cannot capture the direction-distance-oriented local structure, which is essential for 3D object detection. In this work, we propose a Structure-Embedding transFormer (SEFormer), which can not only preserve the local structure as a traditional Transformer but also have the ability to encode the local structure. Compared to the self-attention mechanism in traditional Transformer, SEFormer learns different feature transformations for value points based on the relative directions and distances to the query point. Then we propose a SEFormer-based network for high-performance 3D object detection. Extensive experiments show that the proposed architecture can achieve SOTA results on the Waymo Open Dataset, one of the most significant 3D detection benchmarks for autonomous driving. Specifically, SEFormer achieves 79.02% mAP, which is 1.2% higher than existing works. https://github.com/tdzdog/SEFormer. Xiaoyu Feng, Heming Du, Hehe Fan, Yueqi Duan, Yongpan Liu |
AAAI | 2 |
| 2023 | Object-Goal Visual Navigation via Effective Exploration of Relations Among Historical Navigation StatesabstractObject-goal visual navigation aims at steering an agent toward an object via a series of moving steps. Previous works mainly focus on learning informative visual representations for navigation, but overlook the impacts of navigation states on the effectiveness and efficiency of navigation. We observe that high relevance among navigation states will cause navigation inefficiency or failure for existing methods. In this paper, we present a History-inspired Navigation Policy Learning (HiNL) framework to estimate navigation states effectively by exploring relationships among historical navigation states. In HiNL, we propose a History-aware State Estimation (HaSE) module to alleviate the impacts of dominant historical states on the current state estimation. Meanwhile, HaSE also encourages an agent to be alert to the current observation changes, thus enabling the agent to make valid actions. Furthermore, we design a History-based State Regularization (HbSR) to explicitly suppress the correlation among navigation states in training. As a result, our agent can update states more effectively while reducing the correlations among navigation states. Experiments on the artificial platform AI2-THOR (i.e., iTHOR and RoboTHOR) demonstrate that HiNL significantly outperforms state-of-the-art methods on both Success Rate and SPL in unseen testing environments. Heming Du, Lincheng Li, Zi Huang, Xin Yu 0002 |
CVPR | 1 |
| 2023 | RVD: A Handheld Device-Based Fundus Video Dataset for Retinal Vessel SegmentationabstractRetinal vessel segmentation is generally grounded in image-based datasets collected with bench-top devices. The static images naturally lose the dynamic characteristics of retina fluctuation, resulting in diminished dataset richness, and the usage of bench-top devices further restricts dataset scalability due to its limited accessibility. Considering these limitations, we introduce the first video-based retinal dataset by employing handheld devices for data acquisition. The dataset comprises 635 smartphone-based fundus videos collected from four different clinics, involving 415 patients from 50 to 75 years old. It delivers comprehensive and precise annotations of retinal structures in both spatial and temporal dimensions, aiming to advance the landscape of vasculature segmentation. Specifically, the dataset provides three levels of spatial annotations: binary vessel masks for overall retinal structure delineation, general vein-artery masks for distinguishing the vein and artery, and fine-grained vein-artery masks for further characterizing the granularities of each artery and vein. In addition, the dataset offers temporal annotations that capture the vessel pulsation characteristics, assisting in detecting ocular diseases that require fine-grained recognition of hemodynamic fluctuation. In application, our dataset exhibits a significant domain shift with respect to data captured by bench-top devices, thus posing great challenges to existing methods. Thanks to rich annotations and data scales, our dataset potentially paves the path for more advanced retinal analysis and accurate disease diagnosis. In the experiments, we provide evaluation metrics and benchmark results on our dataset, reflecting both the potential and challenges it offers for vessel segmentation tasks. We hope this challenging dataset would significantly contribute to the development of eye disease diagnosis and early prevention. Wahiduzzaman Khan, Hongwei Sheng, Hu Zhang 0005, Heming Du, Sen Wang 0001, Minas Theodore Coroneo, Farshid Hajati, Sahar Shariflou, Michael Kalloniatis, Jack Phu, Ashish Agar, Zi Huang, S. Mojtaba Golzan, Xin Yu 0002 |
NeurIPS | 4 |
| 2023 | Auslan-Daily: Australian Sign Language Translation for Daily Communication and NewsabstractSign language translation (SLT) aims to convert a continuous sign language video clip into a spoken language. Considering different geographic regions generally have their own native sign languages, it is valuable to establish corresponding SLT datasets to support related communication and research. Auslan, as a sign language specific to Australia, still lacks a dedicated large-scale dataset for SLT.To fill this gap, we curate an Australian Sign Language translation dataset, dubbed Auslan-Daily, which is collected from the Auslan educational TV series and Auslan TV programs. The former involves daily communications among multiple signers in the wild, while the latter comprises sign language videos for up-to-date news, weather forecasts, and documentaries. In particular, Auslan-Daily has two main features: (1) the topics are diverse and signed by multiple signers, and (2) the scenes in our dataset are more complex, e.g., captured in various environments, gesture interference during multi-signers' interactions and various camera positions. With a collection of more than 45 hours of high-quality Auslan video materials, we invite Auslan experts to align different fine-grained visual and language pairs, including video $\leftrightarrow$ fingerspelling, video $\leftrightarrow$ gloss, and video $\leftrightarrow$ sentence. As a result, Auslan-Daily contains multi-grained annotations that can be utilized to accomplish various fundamental sign language tasks, such as signer detection, sign spotting, fingerspelling detection, isolated sign language recognition, sign language translation and alignment. Moreover, we benchmark results with state-of-the-art models for each task in Auslan-Daily. Experiments indicate that Auslan-Daily is a highly challenging SLT dataset, and we hope this dataset will contribute to the development of Auslan and the advancement of sign languages worldwide in a broader context. All datasets and benchmarks are available at Auslan-Daily. Shaozu Yuan, Hongwei Sheng, Heming Du, Xin Yu 0002 |
NeurIPS | 4 |
| 2023 | Weakly-supervised Point Cloud Instance Segmentation with Geometric PriorsabstractThis paper investigates how to leverage more readily acquired annotations, i.e., 3D bounding boxes instead of dense point-wise labels, for instance segmentation. We propose a Weakly-supervised point cloud Instance Segmentation framework with Geometric Priors (WISGP) that allows segmentation models to be trained with 3D bounding boxes of instances. Considering intersections among bounding boxes in a scene would result in ambiguous la- bels, we first group points into two sets, i.e., univocal and equivocal sets, indicating the certainty of a 3D point belonging to an instance, respectively. Specifically, 3D points with clear labels belong to the univocal set while the rest are grouped into the equivocal set. To assign reliable labels to points in the equivocal set, we design a Geometry-guided Label Propagation (GLP) scheme that progressively propagates labels to linked points based on geometric structure, e.g., polygon meshes and superpoints. Afterwards, we train an instance segmentation model with the univocal points and equivocal points labeled by GLP, and then employ it to assign pseudo labels for the remainder of the unlabeled points. Lastly, we retrain the model with all the labeled points to achieve better instance segmentation performance. Experiments on large-scale datasets ScanNet-v2 and S3DIS demonstrate that WISGP is superior to competing weakly-supervised algorithms and even on par with a few fully-supervised ones. Heming Du, Xin Yu 0002, Farookh Khadeer Hussain, Mohammad Ali Armin, Lars Petersson, Weihao Li 0005 |
WACV | 1 |
| 2022 | Monocular Camera-Based Point-Goal Navigation by Learning Depth Channel and Cross-Modality Pyramid FusionabstractFor a monocular camera-based navigation system, if we could effectively explore scene geometric cues from RGB images, the geometry information will significantly facilitate the efficiency of the navigation system. Motivated by this, we propose a highly efficient point-goal navigation framework, dubbed Geo-Nav. In a nutshell, our Geo-Nav consists of two parts: a visual perception part and a navigation part. In the visual perception part, we firstly propose a Self-supervised Depth Estimation network (SDE) specially tailored for the monocular camera-based navigation agent. Our SDE learns a mapping from an RGB input image to its corresponding depth image by exploring scene geometric constraints in a self-consistency manner. Then, in order to achieve a representative visual representation from the RGB inputs and learned depth images, we propose a Cross-modality Pyramid Fusion module (CPF). Concretely, our CPF computes a patch-wise cross-modality correlation between different modal features and exploits the correlation to fuse and enhance features at each scale. Thanks to the patch-wise nature of our CPF, we can fuse feature maps at high resolution, allowing our visual network to perceive more image details. In the navigation part, our extracted visual representations are fed to a navigation policy network to learn how to map the visual representations to agent actions effectively. Extensive experiments on a widely-used multiple-room environment Gibson demonstrate that Geo-Nav outperforms the state-of-the-art in terms of efficiency and effectiveness. Tianqi Tang 0002, Heming Du, Xin Yu 0002, Yi Yang 0001 |
AAAI | 2 |
| 2021 | VTNet: Visual Transformer Network for Object Goal Navigation
Heming Du, Xin Yu 0002, Liang Zheng 0001 |
ICLR | 1 |
| 2020 | Learning Object Relation Graph and Tentative Policy for Visual Navigation
Heming Du, Xin Yu 0002, Liang Zheng 0001 |
ECCV (7) | 1 |
| 2019 | A Neural Micro-Expression RecognizerabstractRecognizing micro-expressions underpins significant and critical research and significant application. We speculate that this problem requires the understanding of the subtle face movement, integration of face structures and a solution of limited training data. In this paper, we build an effective micro-expression recognition system that leverages techniques stemming from these speculations. First, we introduce an optical flow method based on the onset frame and the apex frame to encode the subtle face motion. This has already been validated by prior research. Second, to obtain discriminative representations from the rigid face structures, part-based average pooling is proposed to inject structure priors to the network. Finally, because the system suffers from small training sets, we propose to transfer domain knowledge from macro-expression recognition tasks to micro-expression recognition. Specifically, we adopt two domain adaptation techniques including adversarial training and expression magnification and reduction (EMR). Through experiment, we show that the proposed system achieves very competitive results on the 2ndMicro-Expression Grand Challenge (MEGC). Yuchi Liu, Heming Du, Liang Zheng 0001, Tom Gedeon |
FG | 2 |
| 2019 | Spotting Visual Keywords from Temporal Sliding WindowsabstractVisual Keyword Spotting (KWS), as a newly proposed task deriving from visual speech recognition, has plenty of room for improvements. This paper details our Visual Keyword Spotting system used in the first Mandarin Audio-Visual Speech Recognition Challenge (MAVSR 2019). With the assumption that the vocabularies of target dataset are a subset of the vocabulary of the training set, we proposed a simple and scalable classification based strategy that achieves 19.0% mean average precision (mAP) on this challenge. Our method is based on the idea of using sliding windows to bridge between the word-level dataset and the sentence-level dataset, showing that a strong word level classifier can be directly used in building sentence embedding, thereby making it possible to build a KWS system. Yue Yao 0001, Heming Du, Liang Zheng 0001, Tom Gedeon |
ICMI | 3 |