EDBT 2026 Demo / reviewers in the wild / expert
Cong Wang 0039
dblp:18/2771-39
· DBLP profile ↗
12ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0002-4539-2525ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Lafa: Unlocking Superior Memory Efficiency via Adaptive Metadata Strategy for Scalable Large-Scale Dataset LoadingabstractThe rapid growth of deep learning models and the increasing demand for large-scale datasets have posed unprece dented challenges for data loading and memory management. Existing frameworks (e.g., PyTorch, TensorFlow) often encounter performance bottlenecks when handling large datasets resulting in inefficiencies and excessive memory usage. To address these issues, we propose Lafa, a dynamic metadata loading mechanism optimized for efficient large-scale dataset processing. Lafa introduces the .Lafa format and an adaptive loading strategy with three modes to balance memory usage and loading performance, along with a local shuffle approach that reduces memory overhead and computational complexity while preserving data randomness. Experimental results on GPU (RTX 3090) and Ascend (910A) platforms demonstrate that Lafa significantly improves memory efficiency compared to existing frameworks. Specifically, for every 10 million samples loaded, Lafa reduces additional memory consumption by a factor of 1.33× to 31.34× across various dataset types, relative to the most memory-efficient baseline among PyTorch, TensorFlow, and MindSpore. Cong Wang 0039, Yang Luo 0006, Ke Wang 0065, Hui Zhang 0044, Naijie Gu, Wenzhuo Du, Fan Yu 0004, Jun Yu 0001 |
IEEE Trans. Big Data | 1 |
| 2025 | LVLM-MIR: Large Vision-Language Model with Parameter-Efficient Fine-Tuning for Multimodal Interleaved ReasoningabstractMultimodal interleaved reasoning, which requires models to understand interleaved image-text sequences and multiple images, is a critical challenge in contemporary AI. This paper proposes a parameter-efficient fine-tuning framework based on Large Vision-Language Models, with Qwen2.5-VL as the backbone and Low-Rank Adaptation for task-specific adaptation. The framework integrates four stages: multimodal input preprocessing to align with pre-training distributions, visual feature extraction via a modified Vision Transformer, cross-modal fusion via attention mechanisms, and response generation via an autoregressive decoder. By freezing pre-trained weights and fine-tuning low-rank adapters in both visual and language modules, it balances preserving general multimodal knowledge with optimizing target tasks, achieving high performance with low computational overhead. On the MIRAGE Challenge Track A Dataset, it performs strongly across subtasks, achieving an aggregate score of 0.7857 and securing second place in the challenge. Ablation studies confirm that joint LoRA fine-tuning of visual and language modules yields optimal results; limitations in fine-grained visual difference tasks indicate future directions in enhancing subtle feature capture and adaptive cross-modal alignment. Jun Yu 0001, Xilong Lu, Cong Wang 0039, Qiang Ling 0001 |
ACM Multimedia | 3 |
| 2025 | Breaking barriers in 3D point cloud data processing: A unified system for efficient storage and high-throughput loading
Cong Wang 0039, Yang Luo 0006, Ke Wang 0065, Yanfei Cao, Xiangzhi Tao, Dongjie Geng, Naijie Gu, Jun Yu 0001, Fan Yu 0004, Zhengdong Wang, Shouyang Dong |
Expert Syst. Appl. | 1 |
| 2025 | Faster and Stronger: Unleashing Data Processing Potential Through Hardware HeterogeneityabstractWith the rapid advancement of AI technology, there has been a substantial surge in the need for computational resources. Particularly in deep learning, machine learning, and large-scale data analysis, the processing of extensive datasets necessitates exceptionally high levels of computational efficacy and speed. Conventional homogeneous computing platforms, predominantly reliant on Central Processing Units (CPU), have encountered challenges in meeting the escalating demands for high-performance computing. Consequently, this study advocates for heterogeneous hardware acceleration technology, strategically migrating data operations from CPU to varied hardware components (e.g. GPU, NPU) to enhance processing efficiency and computational performance during the data preprocessing phase. We conducted experiments to evaluate the impact of utilizing hardware heterogeneous acceleration technologies on data processing speed under various workloads and system hardware configurations. By adjusting parameters like batch size and CPU utilization rates, we compared the performance of frameworks that support hardware heterogeneity with popular deep learning frameworks (e.g. PyTorch and TensorFlow) across various hardware configurations and neural network models. Empirical findings demonstrate that the system framework optimized through heterogeneous hardware acceleration technology (the preprocessing speed is improved in all the given experimental environment tests) exhibits commendable universality and superiority in performance. Codes are available at https://github.com/mindspore-ai/mindspore. Cong Wang 0039, Yang Luo 0006, Wenzhuo Du, Ke Wang 0065, Naijie Gu, Jun Yu 0001 |
IEEE Internet Things J. | 1 |
| 2024 | Variational Distillation for Multi-View LearningabstractInformation Bottleneck (IB) provides an information-theoretic principle for multi-view learning by revealing the various components contained in each viewpoint. This highlights the necessity to capture their distinct roles to achieve view-invariance and predictive representations but remains under-explored due to the technical intractability of modeling and organizing innumerable mutual information (MI) terms. Recent studies show that sufficiency and consistency play such key roles in multi-view representation learning, and could be preserved via a variational distillation framework. But when it generalizes to arbitrary viewpoints, such strategy fails as the mutual information terms of consistency become complicated. This paper presents Multi-View Variational Distillation (MV$^{2}$D), tackling the above limitations for generalized multi-view learning. Uniquely, MV$^{2}$D can recognize useful consistent information and prioritize diverse components by their generalization ability. This guides an analytical and scalable solution to achieving both sufficiency and consistency. Additionally, by rigorously reformulating the IB objective, MV$^{2}$D tackles the difficulties in MI optimization and fully realizes the theoretical advantages of the information bottleneck principle. We extensively evaluate our model on diverse tasks to verify its effectiveness, where the considerable gains provide key insights into achieving generalized multi-view representations under a rigorous information-theoretic principle. Zhizhong Zhang 0001, Cong Wang 0039, Wensheng Zhang 0002, Yanyun Qu, Lizhuang Ma, Zongze Wu 0001, Yuan Xie 0006, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Weakly Supervised 3D Segmentation via Receptive-Driven Pseudo Label Consistency and Structural ConsistencyabstractAs manual point-wise label is time and labor-intensive for fully supervised large-scale point cloud semantic segmentation, weakly supervised method is increasingly active. However, existing methods fail to generate high-quality pseudo labels effectively, leading to unsatisfactory results. In this paper, we propose a weakly supervised point cloud semantic segmentation framework via receptive-driven pseudo label consistency and structural consistency to mine potential knowledge. Specifically, we propose three consistency contrains: pseudo label consistency among different scales, semantic structure consistency between intra-class features and class-level relation structure consistency between pair-wise categories. Three consistency constraints are jointly used to effectively prepares and utilizes pseudo labels simultaneously for stable training. Finally, extensive experimental results on three challenging datasets demonstrate that our method significantly outperforms state-of-the-art weakly supervised methods and even achieves comparable performance to the fully supervised methods. Yuxiang Lan, Yachao Zhang 0001, Yanyun Qu, Cong Wang 0039, Yuan Xie 0006, Zongze Wu 0001 |
AAAI | 4 |
| 2023 | Feedback Chain Network for Hippocampus SegmentationabstractThe hippocampus plays a vital role in the diagnosis and treatment of many neurological disorders. Recent years, deep learning technology has made great progress in the field of medical image segmentation, and the performance of related tasks has been constantly refreshed. In this paper, we focus on the hippocampus segmentation task and propose a novel hierarchical feedback chain network. The feedback chain structure unit learns deeper and wider feature representation of each encoder layer through the hierarchical feature aggregation feedback chains, and achieves feature selection and feedback through the feature handover attention module. Then, we embed a global pyramid attention unit between the feature encoder and the decoder to further modify the encoder features, including the pair-wise pyramid attention module for achieving adjacent attention interaction and the global context modeling module for capturing the long-range knowledge. The proposed approach achieves state-of-the-art performance on three publicly available datasets, compared with existing hippocampus segmentation approaches. The code and results can be found from the link of https://github.com/easymoneysniper183/sematic_seg. Heyu Hung, Runmin Cong, Lianhe Yang, Cong Wang 0039, Sam Kwong |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2022 | Task-Level Self-Supervision for Cross-Domain Few-Shot LearningabstractLearning with limited labeled data is a long-standing problem. Among various solutions, episodic training progres-sively classifies a series of few-shot tasks and thereby is as-sumed to be beneficial for improving the model’s generalization ability. However, recent studies show that it is eveninferior to the baseline model when facing domain shift between base and novel classes. To tackle this problem, we pro-pose a domain-independent task-level self-supervised (TL-SS) method for cross-domain few-shot learning.TL-SS strategy promotes the general idea of label-based instance-levelsupervision to task-level self-supervision by augmenting mul-tiple views of tasks. Two regularizations on task consistencyand correlation metric are introduced to remarkably stabi-lize the training process and endow the generalization ability into the prediction model. We also propose a high-order associated encoder (HAE) being adaptive to various tasks.By utilizing 3D convolution module, HAE is able to generate proper parameters and enables the encoder to flexibly toany unseen tasks. Two modules complement each other andshow great promotion against state-of-the-art methods experimentally. Finally, we design a generalized task-agnostic test,where our intriguing findings highlight the need to re-think the generalization ability of existing few-shot approaches. Wang Yuan, Zhizhong Zhang 0001, Cong Wang 0039, Yuan Xie 0006, Lizhuang Ma |
AAAI | 3 |
| 2022 | Optimal Transport for Label-Efficient Visible-Infrared Person Re-Identification
Jiangming Wang, Zhizhong Zhang 0001, Mingang Chen, Cong Wang 0039, Bin Sheng 0001, Yanyun Qu, Yuan Xie 0006 |
ECCV (24) | 5 |
| 2022 | Pyramidal Transformer with Conv-Patchify for Person Re-identificationabstractThe robust and discriminative feature extraction is the key component in person re-identification (Re-ID). The major weakness of conventional convolution neural network (CNN) based methods is that they cannot extract long-range information from diverse parts, which can be alleviated by recently developed Transformers. Existing vision Transformers show their power on various vision tasks. However, they (i) cannot address translation problems and different viewpoints; (ii) cannot capture detailed features to discriminate people with a similar appearance. In this paper, we propose a powerful Re-ID baseline built on top of the pyramidal transformer with conv-patchify operation, termed PTCR, which inherits the advantages of both CNN and Transformer. The pyramidal structure captures multi-scale fine-grained features, while the conv-patchify enhances the robustness against translation. Moreover, we additionally design two novel modules to improve the robust feature learning. A Token Perception module augments the patch embeddings to enhance the robustness against perturbation and viewpoint changes, while the Auxiliary Embedding module integrates the auxiliary information (cam ID, pedestrian attributes, etc) to reduce feature bias caused by non-visual factors. Our method is validated through extensive experiments to show its superior performance with abundant ablation studies. Notably, without re-ranking, we achieve 98.0% Rank-1 on Market-1501 and 88.6% Rank-1 on MSMT17, significantly outperforming the counterparts. The code is available at: https://github.com/lihe404/PTCR He Li 0054, Mang Ye, Cong Wang 0039, Bo Du 0001 |
ACM Multimedia | 3 |
| 2022 | Self-supervised Exclusive Learning for 3D Segmentation with Cross-Modal Unsupervised Domain Adaptationabstract2D-3D unsupervised domain adaptation (UDA) tackles the lack of annotations in a new domain by capitalizing the relationship between 2D and 3D data. Existing methods achieve considerable improvements by performing cross-modality alignment in a modality-agnostic way, failing to exploit modality-specific characteristic for modeling complementarity. In this paper, we present self-supervised exclusive learning for cross-modal semantic segmentation under the UDA scenario, which avoids the prohibitive annotation. Specifically, two self-supervised tasks are designed, named "plane-to-spatial'' and "discrete-to-textured''. The former helps the 2D network branch improve the perception of spatial metrics, and the latter supplements structured texture information for the 3D network branch. In this way, modality-specific exclusive information can be effectively learned, and the complementarity of multi-modality is strengthened, resulting in a robust network to different domains. With the help of the self-supervised tasks supervision, we introduce a mixed domain to enhance the perception of the target domain by mixing the patches of the source and target domain samples. Besides, we propose a domain-category adversarial learning with category-wise discriminators by constructing the category prototypes for learning domain-invariant features. We evaluate our method on various multi-modality domain adaptation settings, where our results significantly outperform both uni-modality and multi-modality state-of-the-art competitors. Yachao Zhang 0001, Miaoyu Li, Yuan Xie 0006, Cuihua Li, Cong Wang 0039, Zhizhong Zhang 0001, Yanyun Qu |
ACM Multimedia | 5 |
| 2022 | Dual Mutual Learning for Cross-Modality Person Re-IdentificationabstractCross-modality person re-identification (Re-ID) is more challenging than traditional visible Re-ID due to the huge cross-modality gap from heterogeneous images. To alleviate this problem, existing methods often utilize a dual path learning framework equipped with metric loss to learn discriminative features. Despite effectiveness, the inevitable degeneration of intra-modality discrimination by taking cross-modality discrimination into consideration is unsolvable. Such degeneration substantially hinders the model’s capability of further improving feature representations. To mitigate this degeneration, we propose a Dual Mutual Learning (DML) method for cross-modality Re-ID which conducts mutual learning between the cross-modality and each of two single modalities. We design a triple-branch deep model containing the RGB and IR branches and the cross-modality branch. The cross-modality branch is designed to learn modality-invariant feature subspace for appearance similarity measurement. Both the RGB branch and IR branch provide attention supervision information to the cross-modality branch for attention feature alignment so as to enhance the intra-modality discrimination. Experimental results on two standard benchmarks demonstrate DML is superior to state-of-the-art methods. Demao Zhang, Zhizhong Zhang 0001, Ying Ju 0002, Cong Wang 0039, Yuan Xie 0006, Yanyun Qu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |