EDBT 2026 Demo / reviewers in the wild / expert
Mingfeng Zha
dblp:310/3729
· DBLP profile ↗
10ranked-venue papers
8as first author
10since 2021 · last 2026
0000-0003-0186-2940ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 7 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Seeing Beyond Illusion: Generalized and Efficient Mirror DetectionabstractReflective imaging enables the mirror imagings and physical entities to possess identical attributes, e.g., color and shape. Current mirror detection (MD) methods primarily rely on designing functional components to establish the correlation and disparities between the imagings and entities, thereby identifying the mirror regions. However, the exploration of extended scenes with dynamic content changes is rarely investigated. Therefore, we propose the MirrorSAM designed for MD based on the Segment Anything Model (SAM). Specifically, due to the varying reflections produced by mirrors in different positions and the complex visual space that interferes with localization, we design the hierarchical mixture of direction experts (HMDE) in the low-rank space to reduce biases towards entities in SAM and dynamically adjust experts based on the input scene. We observe differences in depth between mirrors and adjacent areas, and propose the depth token calibration (DTC), which introduces a learnable depth token to generate the depth map and serve as an error correction factor. We further formulate the selective pixel-prototype contrastive (SPPC) loss, selecting partially confusable samples to promote the decoupling of mirror and non-mirror representations. Extensive experiments conducted on four mirror benchmarks and two settings demonstrate that our approach surpasses state-of-the-art methods with few trainable parameters and FLOPs. We further extend to four transparent surface benchmarks to validate generalization. Mingfeng Zha, Guoqing Wang 0001, Tianyu Li 0003, Wei Dong 0010, Peng Wang 0023, Yang Yang 0002 |
AAAI | 1 |
| 2026 | Think Twice Before Determining: Toward Scene-Aware Visual Reasoning for Mirror DetectionabstractMirror detection (MD) aims to overcome interference caused by reflections and locate mirror regions. Existing methods focus on designing components to explicitly establish the associations between physical entities and corresponding imagings, or utilizing rotation to construct symmetric consistency. We observe that: a) incomplete and incorrect correspondence between entities and imagings; b) other physical materials (e.g., glass) exhibit characteristics partially similar to mirrors, causing confusion when they co-occur; c) complex interfering factors (e.g., occlusion) and reflection mechanisms may expand vector space several times over. To address these issues in a unified manner, we formulate the scene-aware visual reasoning network (SVRNet) based on visual prompts. Specifically, we construct the prototype-guided prompt chain reasoning (PPCR) that generates a mixed chain of thought reasoning based on maximal difference heterogeneous prototypes to construct comprehensive spatial location and semantic perception. Noise may accumulate gradually through the chain, and crucial clues may also disappear. Therefore, we design the prompt evolution (PE) to filter out noise and enhance the coupling between prompts. We further develop the mixture of prompt injection expert (MPIE) to dynamically select the optimal injection strategy in the low-rank space based on specific scene. Due to reflection interference and random parameter space introducing potential ambiguity, we formulate the three-way evidence-aware (TEA) loss to quantify the uncertainty, thereby providing reliable predictions. To leverage historical knowledge and further disentangle representations, we propose the frequency prototype contrastive (FPC) loss for learning more generalizable features across images. Finally, we relabel 25,828 images and formulate the first point-supervised MD framework. Extensive experiments conducted on four mirror benchmarks under three settings demonstrate that our method surpasses state-of-the-art approaches. Promising results are also achieved on six related benchmarks, showing its generality. Mingfeng Zha, Guoqing Wang 0001, Yunqiang Pei, Tianyu Li 0003, Xiongxin Tang, Jiayi Ma 0001, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Hierarchical Consistency Learning for Test-Time Adaptation in Camouflage PerceptionabstractCamouflaged object detection (COD) aims to localize targets that exhibit minimal perceptual differences from backgrounds through physical attributes. Existing methods, constrained by the static train-then-freeze paradigm, suffer from domain rigidity and annotation dependency, limiting their adaptability to scene variations and unseen camouflage patterns. To overcome these, we propose the hierarchical consistency learning (HCL) framework, which integrates test-time adaptation for dynamic representation recalibration. Specifically, we design the hierarchical representation reconstruction (HRR) to alleviate feature entanglement by synergizing spatial reconstruction with dual-stream frequency-domain decomposition, enhancing robustness against appearance homogenization. The pixel and spectrum inference provide structural and contextual priors. We further introduce task affinity guidance (TAG) to propagate knowledge across branches via channel-wise affinity, aligning local discriminative cues and mitigating semantic drift. To ensure semantic invariance, we formulate the prototype consistency calibration (PCC), which aggregates region features into compact prototypes and establishes prototype-feature similarity. This imposes implicit and hierarchical constraints that bridge task and representation gaps. Extensive experiments across four camouflaged and four underwater object benchmarks, under three degradation settings, demonstrate that our method consistently outperforms state-of-the-art approaches, highlighting its robustness and generalization under distribution shifts. Mingfeng Zha, Tianyu Li 0003, Guoqing Wang 0001, Yunqiang Pei, Chaofan Qiao, Jiening Zhang, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Image Process. | 1 |
| 2025 | Implicit Counterfactual Learning for Audio-Visual SegmentationabstractAudio-visual segmentation (AVS) aims to segment objects in videos based on audio cues. Existing AVS methods are primarily designed to enhance interaction efficiency but pay limited attention to modality representation discrepancies and imbalances. To overcome this, we propose the implicit counterfactual framework (ICF) to achieve unbiased cross-modal understanding. Due to the lack of semantics, heterogeneous representations may lead to erroneous matches, especially in complex scenes with ambiguous visual content or interference from multiple audio sources. We introduce the multi-granularity implicit text (MIT) involving video-, segment- and frame-level as the bridge to establish the modality-shared space, reducing modality gaps and providing prior guidance. Visual content carries more information and typically dominates, thereby marginalizing audio features in the decision-making. To mitigate knowledge preference, we propose the semantic counterfactual (SC) to learn orthogonal representations in the latent space, generating diverse counterfactual samples, thus avoiding biases introduced by complex functional designs and explicit modifications of text structures or attributes. We further formulate the collaborative distribution-aware contrastive learning (CDCL), incorporating factual-counterfactual and inter-modality contrasts to align representations, promoting cohesion and decoupling. Extensive experiments on three public datasets validate that the proposed method achieves state-of-the-art performance. Mingfeng Zha, Tianyu Li 0003, Guoyin Wang 0001, Peng Wang 0023, Yang Yang 0002, Heng Tao Shen |
ICCV | 1 |
| 2025 | AttentionAR: AR Adaptation and Warning for Real-World Safety via Attention Modeling and MLLM Reasoning
Yunqiang Pei, Renming Huang, Mingfeng Zha, Guoqing Wang 0001, Peng Wang 0023, Qiao Kang, Yang Yang 0002, Heng Tao Shen |
UIST | 3 |
| 2025 | Heterogeneous Experts and Hierarchical Perception for Underwater Salient Object DetectionabstractExisting underwater salient object detection (USOD) methods design fusion strategies to integrate multimodal information, but lack exploration of modal characteristics. To address this, we separately leverage the RGB and depth branches to learn disentangled representations, formulating the heterogeneous experts and hierarchical perception network (HEHP). Specifically, to reduce modal discrepancies, we propose the hierarchical prototype guided interaction (HPI), which achieves fine-grained alignment guided by the semantic prototypes, and then refines with complementary modalities. We further design the mixture of frequency experts (MoFE), where experts focus on modeling high- and low-frequency respectively, collaborating to explicitly obtain hierarchical representations. To efficiently integrate diverse spatial and frequency information, we formulate the four-way fusion experts (FFE), which dynamically selects optimal experts for fusion while being sensitive to scale and orientation. Since depth maps with poor quality inevitably introduce noises, we design the uncertainty injection (UI) to explore high uncertainty regions by establishing pixel-level probability distributions. We further formulate the holistic prototype contrastive (HPC) loss based on semantics and patches to learn compact and general representations across modalities and images. Finally, we employ varying supervision based on branch distinctions to implicitly construct difference modeling. Extensive experiments on two USOD datasets and four relevant underwater scene benchmarks validate the effect of the proposed method, surpassing state-of-the-art binary detection models. Impressive results on seven natural scene benchmarks further demonstrate the scalability. Mingfeng Zha, Guoqing Wang 0001, Yunqiang Pei, Tianyu Li 0003, Xiongxin Tang, Chongyi Li, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Image Process. | 1 |
| 2024 | Weakly-Supervised Mirror Detection via Scribble AnnotationsabstractMirror detection is of great significance for avoiding false recognition of reflected objects in computer vision tasks. Existing mirror detection frameworks usually follow a supervised setting, which relies heavily on high quality labels and suffers from poor generalization. To resolve this, we instead propose the first weakly-supervised mirror detection framework and also provide the first scribble-based mirror dataset. Specifically, we relabel 10,158 images, most of which have a labeled pixel ratio of less than 0.01 and take only about 8 seconds to label. Considering that the mirror regions usually show great scale variation, and also irregular and occluded, thus leading to issues of incomplete or over detection, we propose a local-global feature enhancement (LGFE) module to fully capture the context and details. Moreover, it is difficult to obtain basic mirror structure using scribble annotation, and the distinction between foreground (mirror) and background (non-mirror) features is not emphasized caused by mirror reflections. Therefore, we propose a foreground-aware mask attention (FAMA), integrating mirror edges and semantic features to complete mirror regions and suppressing the influence of backgrounds. Finally, to improve the robustness of the network, we propose a prototype contrast loss (PCL) to learn more general foreground features across images. Extensive experiments show that our network outperforms relevant state-of-the-art weakly supervised methods, and even some fully supervised methods. The dataset and codes are available at https://github.com/winter-flow/WSMD. Mingfeng Zha, Yunqiang Pei, Guoqing Wang 0001, Tianyu Li 0003, Yang Yang 0002, Wenbin Qian, Heng Tao Shen |
AAAI | 1 |
| 2024 | Emotion Recognition in HMDs: A Multi-task Approach Using Physiological Signals and Occluded FacesabstractPrior research on emotion recognition in extended reality (XR) has faced challenges due to the occlusion of facial expressions by Head-Mounted Displays (HMDs). This limitation hinders accurate Facial Expression Recognition (FER), which is crucial for immersive user experiences. This study aims to overcome the occlusion challenge by integrating physiological signals with partially visible facial expressions to enhance emotion recognition in XR environments. We employed a multi-task approach, utilizing a feature-level fusion to fuse Electroencephalography (EEG) and Galvanic Skin Response (GSR) signals with occluded facial expressions. The model predicts valence and arousal simultaneously from both macro-and micro-expression. Our method demonstrated improved accuracy in emotion recognition under partial occlusion conditions. The integration of temporal physiological signals with other modalities significantly enhanced performance, particularly for half-face emotion recognition. The study presents a novel approach to emotion recognition in XR, addressing the limitations of facial occlusion by HMDs. The findings suggest that physiological signals are vital for interpreting emotions in occluded scenarios, offering potential for real-time applications and advancing social XR applications. Yunqiang Pei, Jialei Tang, Qihang Tang, Mingfeng Zha, Dongyu Xie, Guoqing Wang 0001, Zhitao Liu, Ning Xie 0003, Peng Wang 0023, Yang Yang 0002, Heng Tao Shen |
ACM Multimedia | 4 |
| 2024 | Dual Domain Perception and Progressive Refinement for Mirror DetectionabstractMirror detection aims to discover mirror regions in images to avoid misidentifying reflected objects. Existing methods mainly mine clues from spatial domain. We observe that the frequencies inside and outside the mirror region are distinctive. Besides, the low-frequency representing the feature semantics can help to locate the mirror region, and the high-frequency representing the details can refine it. Motivated by this, we introduce frequency guidance and propose the dual domain perception progressive refinement network (DPRNet) to mine dual-domain information. Specifically, we first decouple the images into high-frequency and low-frequency components by Laplace pyramid and vision Transformer, respectively, and design the frequency interaction alignment (FIA) module to integrate frequency features to initially localize the mirror region. To handle scale variations, we propose the multi-order feature perception (MOFP) module to adaptively aggregate adjacent features with progressive and gating mechanisms. We further propose the separation-based difference fusion (SDF) module to establish associations between entities and imagings and discover the correct boundary to mine the complete mirror region. Extensive experiments show that DPRNet outperforms the state-of-the-art method by an average of 3% with only about one-fifth of the parameters and FLOPs on four datasets. Our DPRNet also achieves promising performance on remote sensing and camouflage scenarios, validating its generalization. The code is available athttps://github.com/winter-flow/DPRNet. Mingfeng Zha, Feiyang Fu, Yunqiang Pei, Guoqing Wang 0001, Tianyu Li 0003, Xiongxin Tang, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Multifeature Transformation and Fusion-Based Ship Detection With Small Targets and Complex BackgroundsabstractWith the development of deep learning, synthetic aperture radar (SAR) image ship detection based on the convolutional neural network has made significant progress. However, there are two problems. 1) The false alarm detection rate is high due to complex background and coherent speckle noise interference. 2) For smaller ship targets, missed detection is prone to occur. In this letter, a novel ship detection model (MFTF-Net) based on multi-feature transformation and fusion is proposed to address the issues. First, to avoid the randomness of initial point selection and the influence of outlier points, the anchor frame clustering approach based on the K-medians++ algorithm is presented to cluster the object candidate frames. Second, the low-level feature information is passed to the high level by constructing a local enhancement network; then, an improved Transformer structure is introduced to replace the last convolutional block of the backbone network to obtain rich contextual information. Finally, a four-scale residual feature fusion network is designed, which fully fuses the object’s detailed and semantic information. In addition, improved convolutional block attention module (CBAM) and squeeze and excitation (SE) attention mechanisms are applied in the lower two layers and upper two layers of the network output to reduce the interference of confusing information, respectively. The experimental results demonstrate that the proposed method is superior to the state-of-the-art thirteen baseline models on SAR ship detection dataset (SSDD), high-resolution SAR images dataset (HRSID), and SAR-ship-dataset public datasets in terms of the mAP, recall, accuracy, and F1 metrics. Mingfeng Zha, Wenbin Qian, Wenji Yang, Yilu Xu |
IEEE Geosci. Remote. Sens. Lett. | 1 |