Heng Cai

dblp:68/7829 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Segmentation and scene understanding · 54% Vision and language · 41% Generative modeling · 5%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Segmentation and scene understanding
annotation-efficient segmentation
0.712023
Orthogonal Annotation Benefits Barely-supervised Medical Image Segmentation · CVPR 2023
Computer vision › Vision and language
image-text retrieval
0.712023
CCMB: A Large-scale Chinese Cross-modal Benchmark · ACM Multimedia 2023
Computer vision › Segmentation and scene understanding
medical image segmentation
0.712023
Orthogonal Annotation Benefits Barely-supervised Medical Image Segmentation · CVPR 2023
Computer vision › Segmentation and scene understanding › medical image segmentation
semi-supervised segmentation
0.712023
Orthogonal Annotation Benefits Barely-supervised Medical Image Segmentation · CVPR 2023
Computer vision › Vision and language
vision-language pretraining
0.712023
CCMB: A Large-scale Chinese Cross-modal Benchmark · ACM Multimedia 2023
Computer vision › Vision and language
image captioning
0.212023
CCMB: A Large-scale Chinese Cross-modal Benchmark · ACM Multimedia 2023
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.212023
CCMB: A Large-scale Chinese Cross-modal Benchmark · ACM Multimedia 2023

Methods — techniques the papers use, named apart from their topics

registration-based pseudo labeling · 0.7pre-ranking and ranking · 0.7orthogonal annotation · 0.7knowledge distillation · 0.7dual-network co-training · 0.7
YearPublicationVenuePosition
2026 Lightmap Compression with Color-Coherent UV Clustering and Cascade Texture Optimization
abstract
Abstract To address the storage overhead of lightmaps and the limitations of existing compression techniques, we propose a novel UV‐space compression framework based on per‐triangle processing. By mapping triangles to a standardized domain, we cluster and repack color‐coherent regions into a compact atlas, generating a cascade texture refined via differentiable rendering. Experimental results show an average storage reduction of 83% with approximately 10 dB higher PSNR than existing methods. Our approach is the first dedicated lightmap compression framework compatible with standard block‐based formats, offering an effective solution for memory‐efficient 3D asset delivery.
Dehan Chen, Hongyu Huang 0001, Yuzhe Luo, Hao Xu 0049, Yuqing Zhang 0005, Sipeng Yang, Xifeng Gao, Heng Cai, Xiaogang Jin 0001
Comput. Graph. Forum8
2023 Orthogonal Annotation Benefits Barely-supervised Medical Image Segmentation
abstract
Recent trends in semi-supervised learning have significantly boosted the performance of 3D semi-supervised medical image segmentation. Compared with 2D images, 3D medical volumes involve information from different directions, e.g., transverse, sagittal, and coronal planes, so as to naturally provide complementary views. These complementary views and the intrinsic similarity among adjacent 3D slices inspire us to develop a novel annotation way and its corresponding semi-supervised model for effective segmentation. Specifically, we firstly propose the orthogonal annotation by only labeling two orthogonal slices in a labeled volume, which significantly relieves the burden of annotation. Then, we perform registration to obtain the initial pseudo labels for sparsely labeled volumes. Subsequently, by introducing unlabeled volumes, we propose a dual-network paradigm named Dense-Sparse Co-training (DeSCO) that exploits dense pseudo labels in early stage and sparse labels in later stage and meanwhile forces consistent output of two networks. Experimental results on three benchmark datasets validated our effectiveness in performance and efficiency in annotation. For example, with only 10 annotated slices, our method reaches a Dice up to 86.93% on KiTS19 dataset. Our code and models are available at https://github.com/HengCai-NJU/DeSCO.
Heng Cai, Shumeng Li, Lei Qi 0001, Qian Yu 0007, Yinghuan Shi, Yang Gao 0001
CVPR1
2023 3D Medical Image Segmentation with Sparse Annotation via Cross-Teaching Between 3D and 2D Networks
Heng Cai, Lei Qi 0001, Qian Yu 0007, Yinghuan Shi, Yang Gao 0001
MICCAI (3)1
2023 CCMB: A Large-scale Chinese Cross-modal Benchmark
abstract
Vision-language pre-training (VLP) on large-scale datasets has shown premier performance on various downstream tasks. In contrast to plenty of available benchmarks with English corpus, large-scale pre-training datasets and downstream datasets with Chinese corpus remain largely unexplored. In this work, we build a large-scale high-quality Chinese Cross-Modal Benchmark named CCMB for the research community, which contains the currently largest public pre-training dataset Zero and five human-annotated fine-tuning datasets for downstream tasks. Zero contains 250 million images paired with 750 million text descriptions, plus two of the five fine-tuning datasets are also currently the largest ones for Chinese cross-modal downstream tasks. Along with the CCMB, we also develop a VLP framework named R2D2, applying a pre-Ranking + Ranking strategy to learn powerful vision-language representations and a two-way distillation method (i.e., target-guided Distillation and feature-guided Distillation) to further enhance the learning capability. With the Zero and the R2D2 VLP framework, we achieve state-of-the-art performance on twelve downstream datasets from five broad categories of tasks including image-text retrieval, image-text matching, image caption, text-to-image generation, and zero-shot image classification. The datasets, models, and codes are available at https://github.com/yuxie11/R2D2
Chunyu Xie, Heng Cai, Jincheng Li 0002, Fanjing Kong, Jianfei Song, Henrique Morimitsu, Lin Yao 0003, Xiangzheng Zhang, Dawei Leng, Baochang Zhang 0001, Xiangyang Ji, Yafeng Deng
ACM Multimedia2
2023 PLN: Parasitic-Like Network for Barely Supervised Medical Image Segmentation
abstract
It is known that annotations for 3D medical image segmentation tasks are laborious, time-consuming and expensive. Considering the similarities existing in inter-slice and inter-volume, we believe that the delineation way and the model architecture should be tightly coupled. In this paper, by introducing an extremely sparse annotation way of labeling only one slice per 3D image, we investigate a novel barely-supervised segmentation setting with only a few sparsely-labeled images along with a large amount of unlabeled images. To achieve this goal, we present a new parasitic-like network including a registration module (as host) and a semi-supervised segmentation module (as parasite) to deal with inter-slice label propagation and inter-volume segmentation prediction, respectively. Specifically, our parasitism mechanism effectively achieves the collaboration of these two modules through three stages of infection, development and eclosion, providing accurate pseudo-labels for training. Extensive results demonstrate that our framework is capable of achieving high performance on extremely sparse annotation tasks, e.g., we achieve Dice of 84.83% on LA dataset with only 16 labeled slices. The code is available athttps://github.com/ShumengLI/PLN.
Shumeng Li, Heng Cai, Lei Qi 0001, Qian Yu 0007, Yinghuan Shi, Yang Gao 0001
IEEE Trans. Medical Imaging2