EDBT 2026 Demo / reviewers in the wild / expert
He Tang 0002
dblp:37/9553-2
· DBLP profile ↗
19ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0002-8454-1407ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 2 first-author · 13 since 2021Artificial intelligence and machine learning · 8 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MotionCharacter: Fine-Grained Motion Controllable Human Video GenerationabstractRecent advancements in personalized Text-to-Video (T2V) generation have made significant strides in synthesizing character-specific content. However, these methods face a critical limitation: the inability to perform fine-grained control over motion intensity. This limitation stems from an inherent entanglement of action semantics and their corresponding magnitudes within coarse textual descriptions, hindering the generation of nuanced human videos and limiting their applicability in scenarios demanding high precision, such as animating virtual avatars or synthesizing subtle micro-expressions. Furthermore, existing approaches often struggle to preserve high identity fidelity when other attributes are modified. To address these challenges, we introduce MotionCharacter, a framework for high-fidelity human video generation with precise motion control. At its core, MotionCharacter explicitly decouples motion into two independently controllable components: action type and motion intensity. This is achieved through two key technical contributions: (1) a Motion Control Module that leverages textual phrases to specify the action type and a quantifiable metric derived from optical flow to modulate its intensity, guided by a region-aware loss that localizes motion to relevant subject areas; and (2) an ID Content Insertion Module coupled with an ID-Consistency loss to ensure robust identity preservation during dynamic motions. To facilitate training for such fine-grained control, we also curate Human-Motion, a new large-scale dataset with detailed annotations for both motion and facial features. Extensive experiments demonstrate that MotionCharacter achieves substantial improvements over existing methods. Our framework excels in generating videos that are not only identity-consistent but also precisely adhere to specified motion types and intensities. Haopeng Fang, Di Qiu, Binjie Mao, He Tang 0002 |
AAAI | 4 |
| 2026 | OwlCap: Harmonizing Motion-Detail for Video Captioning via HMD-270K and Caption Set Equivalence RewardabstractVideo captioning aims to generate comprehensive and coherent descriptions of the video content, contributing to the advancement of both video understanding and generation. However, existing methods often suffer from motion-detail imbalance, as models tend to overemphasize one aspect while neglecting the other. This imbalance results in incomplete captions, which in turn leads to a lack of consistency in video understanding and generation. To address this issue, we propose solutions from two aspects: 1) Data aspect: We constructed the Harmonizing Motion-Detail 270K (HMD-270K) dataset through a two-stage pipeline: Motion-Detail Fusion (MDF) and Fine-Grained Examination (FGE). 2) Optimization aspect: We introduce the Caption Set Equivalence Reward (CSER) based on Group Relative Policy Optimization (GRPO). CSER enhances completeness and accuracy in capturing both motion and details through unit-to-set matching and bidirectional validation. Based on the HMD-270K supervised fine-tuning and GRPO post-training with CSER, we developed OwlCap, a powerful video captioning Multi-modal Large Language Model (MLLM) with motion-detail balance. Experimental results demonstrate that OwlCap achieves significant improvements compared to baseline models on two benchmarks: the detail-focused VDC (+4.2 Acc) and the motion-focused DREAM-1K (+4.6 F1). Chunlin Zhong, Qiuxia Hou, Zhangjun Zhou, Shuang Hao 0015, Haonan Lu, He Tang 0002, Xiang Bai |
AAAI | 7 |
| 2026 | Toward Universal Instance Shadow Detection Based on Pairwise Grouping With Contrastive Morphological Alignment
Haopeng Fang, Wenfeng Han, He Tang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Rethinking Detecting Salient and Camouflaged Objects in Unconstrained Scenes
Zhangjun Zhou, Chunlin Zhong, Jianuo Huang, Jialun Pei, He Tang 0002 |
ICCV | 7 |
| 2025 | PathVG: A New Benchmark and Dataset for Pathology Visual Grounding
Chunlin Zhong, Shuang Hao 0015, Xiaona Chang, Jiwei Jiang, Xiu Nie, He Tang 0002, Xiang Bai |
MICCAI (13) | 7 |
| 2024 | CoLA: Conditional Dropout and Language-Driven Robust Dual-Modal Salient Object Detection
Shuang Hao 0015, Chunlin Zhong, He Tang 0002 |
ECCV (15) | 3 |
| 2024 | Tri-Modal Confluence with Temporal Dynamics for Scene Graph Generation in Operating Rooms
Diandian Guo, Manxi Lin, Jialun Pei, He Tang 0002, Yueming Jin, Pheng-Ann Heng |
MICCAI (6) | 4 |
| 2024 | CalibNet: Dual-Branch Cross-Modal Calibration for RGB-D Salient Instance SegmentationabstractIn this study, we propose a novel approach for RGB-D salient instance segmentation using a dual-branch cross-modal feature calibration architecture called CalibNet. Our method simultaneously calibrates depth and RGB features in the kernel and mask branches to generate instance-aware kernels and mask features. CalibNet consists of three simple modules, a dynamic interactive kernel (DIK) and a weight-sharing fusion (WSF), which work together to generate effective instance-aware kernels and integrate cross-modal features. To improve the quality of depth features, we incorporate a depth similarity assessment (DSA) module prior to DIK and WSF. In addition, we further contribute a new DSIS dataset, which contains 1,940 images with elaborate instance-level annotations. Extensive experiments on three challenging benchmarks show that CalibNet yields a promising result, i.e., 58.0% AP with 320×480 input size on the COME15K-E test set, which significantly surpasses the alternative frameworks. Our code and dataset will be publicly available at: https://github.com/PJLallen/CalibNet. Jialun Pei, Tao Jiang 0002, He Tang 0002, Nian Liu 0002, Yueming Jin, Deng-Ping Fan, Pheng-Ann Heng |
IEEE Trans. Image Process. | 3 |
| 2023 | Unite-Divide-Unite: Joint Boosting Trunk and Structure for High-accuracy Dichotomous Image SegmentationabstractHigh-accuracy Dichotomous Image Segmentation (DIS) aims to pinpoint category-agnostic foreground objects from natural scenes. The main challenge for DIS involves identifying the highly accurate dominant area while rendering detailed object structure. However, directly using a general encoder-decoder architecture may result in an oversupply of high-level features and neglect the shallow spatial information necessary for partitioning meticulous structures. To fill this gap, we introduce a novel Unite-Divide-Unite Network (UDUN) that restructures and bipartitely arranges complementary features to simultaneously boost the effectiveness of trunk and structure identification. The proposed UDUN proceeds from several strengths. First, a dual-size input feeds into the shared backbone to produce more holistic and detailed features while keeping the model lightweight. Second, a simple Divide-and-Conquer Module (DCM) is proposed to decouple multiscale low- and high-level features into our structure decoder and trunk decoder to obtain structure and trunk information respectively. Moreover, we design a Trunk-Structure Aggregation module (TSA) in our union decoder that performs cascade integration for uniform high-accuracy segmentation. As a result, UDUN performs favorably against state-of-the-art competitors in all six evaluation metrics on overall DIS-TE, i.e., achieving 0.772 weighted F-measure and 977 HCE. Using 1024X1024 input, our model enables real-time inference at 65.3 fps with ResNet-18. The source code is available at https://github.com/PJLallen/UDUN. Jialun Pei, Zhangjun Zhou, Yueming Jin, He Tang 0002, Pheng-Ann Heng |
ACM Multimedia | 4 |
| 2023 | Partitioned Saliency Ranking with Dense Pyramid TransformersabstractIn recent years, saliency ranking has emerged as a challenging task focusing on assessing the degree of saliency at instance-level. Being subjective, even humans struggle to identify the precise order of all salient instances. Previous approaches undertake the saliency ranking by directly sorting the rank scores of salient instances, which have not explicitly resolved the inherent ambiguities. To overcome this limitation, we propose the ranking by partition paradigm, which segments unordered salient instances into partitions and then ranks them based on the correlations among these partitions. The ranking by partition paradigm alleviates ranking ambiguities in a general sense, as it consistently improves the performance of other saliency ranking models. Additionally, we introduce the Dense Pyramid Transformer (DPT) to enable global cross-scale interactions, which significantly enhances feature interactions with reduced computational burden. Extensive experiments demonstrate that our approach outperforms all existing methods. The code for our method is available at https://github.com/ssecv/PSR. Chengxiao Sun, Jialun Pei, Haopeng Fang, He Tang 0002 |
ACM Multimedia | 5 |
| 2023 | FGO-Net: Feature and Gaussian Optimization Network for visual saliency prediction
Jialun Pei, He Tang 0002, Chao Liu 0063, Chuanbo Chen |
Appl. Intell. | 3 |
| 2023 | Depth-Induced Gap-Reducing Network for RGB-D Salient Object Detection: An Interaction, Guidance and Refinement ApproachabstractDepth provides complementary information for salient object detection (SOD). However, the performance of RGB-D SOD methods is usually hindered by low quality depth map, semantic gap cross-modality and intrinsic gap between multi-level features. Although recent RGB-D SOD methods have been embedded into depth quality assessment, these methods do not consider the inconsistency of the depth format across datasets. In this paper, we propose an interpretable and effective mechanism called interference degree (ID) to assess depth quality and reweight the contribution of single-modality features without extra annotation. Then, a cross-modality interaction block (CMIB) is designed to reduce the semantic gap between RGB and depth features with the help of ID mechanism, and a mutually guided cross-level fusion (MGCF) module is designed to reduce the intrinsic gap among multi-level features. Finally, a refinement branch is proposed to enhance the salient regions and suppress the non-salient regions of fused features. Extensive experiments on six benchmark datasets show that the proposed depth-induced gap-reducing network (DIGR-Net) outperforms 20 recent state-of-the-art methods. Jialun Pei, He Tang 0002, Zehua Lyu, Chuanbo Chen |
IEEE Trans. Multim. | 4 |
| 2023 | Transformer-Based Efficient Salient Instance Segmentation Networks With Orientative QueryabstractSalient instance segmentation (SIS) can be considered as the next generation task for the saliency detection community. Most of the existing state-of-the-art methods used for this novel challenging task are built on the mainstream Mask R-CNN architecture. However, this mechanism relies heavily on hand-designed anchors and NMS post-processing. In this paper, we provide a one stage SIS framework with transformers, termed Orientative Query Transformer (OQTR). To leverage the long-range dependencies of transformers, a cross fusion module is designed to efficiently fuse the global features in the encoder and salient query features for salient mask prediction. Furthermore, derived from the center prior in traditional saliency models, we propose an orientative query that is considered as the initial salient object query to accelerate convergence. In addition, to mitigate the issue of the lack of a large-scale dataset with salient instance labels, we collect a new SIS dataset (SIS10 K) containing over 10 K images elaborately annotated with both object- and instance-level labels to promote the community. Without any post-processing, our end-to-end OQTR framework significantly surpasses the top-1 RDPNet by an average of 13.1% AP scores across all three challenging datasets, demonstrating the strong performance of the proposed OQTR. The code and the dataset proposed in this work are available at:https://github.com/ssecv/OQTR. Jialun Pei, Tianyang Cheng, He Tang 0002, Chuanbo Chen |
IEEE Trans. Multim. | 3 |
| 2022 | OSFormer: One-Stage Camouflaged Instance Segmentation with Transformers
Jialun Pei, Tianyang Cheng, Deng-Ping Fan, He Tang 0002, Chuanbo Chen, Luc Van Gool |
ECCV (18) | 4 |
| 2022 | Salient instance segmentation with region and box-level annotations
Jialun Pei, He Tang 0002, Tianyang Cheng, Chuanbo Chen |
Neurocomputing | 2 |
| 2020 | Salient instance segmentation via subitizing and clustering
Jialun Pei, He Tang 0002, Chao Liu 0063, Chuanbo Chen |
Neurocomputing | 2 |
| 2018 | Saliency detection from one time sampling for eye fixation prediction
He Tang 0002, Chuanbo Chen, Xiaobing Pei |
Multim. Tools Appl. | 1 |
| 2017 | A novel local derivative quantized binary pattern for object recognition
Jun Shang, Chuanbo Chen, Xiaobing Pei, Hu Liang, He Tang 0002, Mudar Sarem |
Vis. Comput. | 5 |
| 2016 | Visual Saliency Detection via Sparse Residual and Outlier DetectionabstractThis letter proposes a bottom-up saliency model to predict eye fixation locations. Unlike traditional models that measure saliency by computing local or global distinctness, the proposed model considers saliency as the prediction error, because we believe that image patches or pixels with higher prediction error are more salient than others. The prediction error consists of both mispredicted error and unpredicted error. We propose a new algorithm called sparse residual to compute the mispredicted error. We then adopt outlier detection to compute the unpredicted error. Finally, we obtain the saliency map from merging the two results together via a guided filter. Extensive experiments on three benchmark databases show that our model is superior to 12 state-of-the-art models. He Tang 0002, Chuanbo Chen, Xiaobing Pei |
IEEE Signal Process. Lett. | 1 |