VLDB 2026 Research / reviewers in the wild / expert
Junhui Chen
dblp:149/0983
· DBLP profile ↗
12ranked-venue papers
0as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rejoining Precious Artifacts: Efficiently Bone Stick Rejoining Based Massive Fragment Images by Contour, Script, and TextureabstractRejoining fragment images of precious artifacts is a meaningful task because complete artifacts could provide valuable clues for the research of human civilization. However, existing rejoining methods face several challenges including time-consuming manual annotation, insufficient rejoining accuracy, and prohibitive computation cost. For rejoining fragment images of bone sticks (a precious artifact), we propose a lightweight vision graph neural network called RejoinViG to address these challenges. First, our method avoids time-consuming manual annotation of ballast contour data by experts. Specifically, our method directly takes a pair of fragment images as input and then determines whether the image pair is rejoinable. Second, our method improves rejoining accuracy by contour, script, and texture through dynamically constructing local and global graphs. Third, our method improves rejoining accuracy while reducing computation cost by introducing a new attention mechanism named node self-attention. Extensive experiments demonstrate that our method outperforms the state-of-the-art methods significantly. For example, the Top-1 accuracy of our method is 3.9 times that of SFF-Siam. Surprisingly, our method successfully rejoins a pair of previously unknown but rejoinable fragment images of bone sticks in a real-world scenario. Xingyi Wang, Wen Huang 0002, Mengqiang Hu, Junhui Chen, Weixin Zhao, Wenzheng Xu, Jian Peng 0002 |
AAAI | 4 |
| 2026 | Real-time sparse signal reconstruction via KKT-conditions-driven analog circuit solver
Xing He 0001, Meng Zhang 0030, Tingwen Huang, Junhui Chen, Ruoxi Yu |
Neural Networks | 6 |
| 2026 | F2M: Improving Skin Disease Recognition by Fusing Multi-Source and Multi-Scale Image FeaturesabstractSkin diseases are one of the most common diseases worldwide, and the mismatch between skin disease patients and dermatologists leads to a huge waste of healthcare resources. Accurately matching skin disease patients to appropriate dermatologists by an image-based method of skin disease recognition can reduce the waste of healthcare resources. However, existing image-based methods of skin disease recognition do not fully utilize multi-source and multi-scale features, which leaves us the chance to improve skin disease recognition further. In this paper, we propose a fusion method of multi-source image features and multi-scale image features to improve skin disease recognition. First, we design a fusion module of multi-source image features to integrate multi-source image information. By dual Convolutional Block Attention Module (CBAM) blocks, the fusion module of multi-source image features enhances the feature representation of key regions and then obtains a comprehensive representation of skin diseases. Second, we propose a fusion module of multi-scale image features. By two parallel backbone networks, the fusion module of multi-scale image features can extract deep feature representations from different scales and exploit their complementarity. To validate the effectiveness of our method, we conduct extensive experiments. The experiment results demonstrate that our method outperforms the state-of-the-art method, achieving improvements of 6.30%, 12.52%, 10.85%, 12.16%, and 5.06% in accuracy, precision, recall, F1-score, and AUC, respectively. Xingyi Wang, Wen Huang 0002, Liaoyaqi Wang, Junhui Chen, Jian Peng 0002, Yuping Ran, Xin Ran |
IEEE Trans. Multim. | 5 |
| 2025 | OracleProtoPNet: Oracle Character Recognition with Interpretability
Wen Huang 0002, Junhui Chen, Xingyi Wang, Jian Peng 0002 |
ICDAR (4) | 3 |
| 2025 | From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3DabstractRecent advances in LVLMs have improved vision-language understanding, but they still struggle with spatial perception, limiting their ability to reason about complex 3D scenes. Unlike previous approaches that incorporate 3D representations into models to improve spatial understanding, we aim to unlock the potential of VLMs by leveraging spatially relevant image data. To this end, we introduce a novel 2D spatial data generation and annotation pipeline built upon scene data with 3D ground-truth. This pipeline enables the creation of a diverse set of spatial tasks, ranging from basic perception tasks to more complex reasoning tasks. Leveraging this pipeline, we construct SPAR-7M, a large-scale dataset generated from thousands of scenes across multiple public datasets. In addition, we introduce SPAR-Bench, a benchmark designed to offer a more comprehensive evaluation of spatial capabilities compared to existing spatial benchmarks, supporting both single-view and multi-view inputs. Training on both SPAR-7M and large-scale 2D datasets enables our models to achieve state-of-the-art performance on 2D spatial benchmarks. Further fine-tuning on 3D task-specific datasets yields competitive results, underscoring the effectiveness of our dataset in enhancing spatial reasoning. Yurui Chen, Yueming Xu, Ze Huang, Jilin Mei, Junhui Chen, Yanpeng Zhou, Yu-Jie Yuan, Xinyue Cai, Xingyue Quan, Hang Xu 0004, Li Zhang 0040 |
NeurIPS | 6 |
| 2025 | Neural Scoring: A Refreshed End-to-End Approach for Speaker Verification in Complex ConditionsabstractModern speaker verification systems primarily rely on speaker embeddings, followed by verification based on cosine similarity between the embedding vectors of the enrollment and test utterances. While effective, these methods struggle with multi-talker speech due to the unidentifiability of embedding vectors. In this paper, we propose Neural Scoring, a refreshed end-to-end framework that directly estimates verification posterior probabilities without relying on test-side embeddings, making it more robust to complex conditions, e.g., with multiple talkers. To make the training of such an end-to-end model more efficient, we introduce a large-scale trial e2e training strategy, where each test utterance pairs with a set of enrolled speakers, thus enabling processing of large-scale verification trials per batch. Experiments on VoxCeleb dataset demonstrate that Neural Scoring consistently outperforms both the baseline and competitive methods across various conditions, achieving an overall 70.36% reduction in Equal Error Rate compared to the baseline. Wan Lin, Junhui Chen, Lantian Li, Dong Wang 0013 |
IEEE Signal Process. Lett. | 2 |
| 2024 | An Investigation of Distribution Alignment in Multi-Genre Speaker RecognitionabstractMulti-genre speaker recognition is becoming increasingly popular due to its ability to better represent the complexities of real-world applications. However, a major challenge is the significant shift in the distribution of speaker vectors across different genres. While distribution alignment is a common approach to address this challenge, previous studies have mainly focused on aligning a source domain with a target domain, and the performance of multi-genre data is unknown.This paper presents a comprehensive study of mainstream distribution alignment methods on multi-genre data, where multiple distributions need to be aligned. We analyze various methods both qualitatively and quantitatively. Our experiments on the CN-Celeb dataset show that within-between distribution alignment (WBDA) performs relatively better. However, we also found that none of the investigated methods consistently improved performance in all test cases. This suggests that solely aligning the distributions of speaker vectors may not fully address the challenges posed by multi-genre speaker recognition. Further investigation is necessary to develop a more comprehensive solution. Junhui Chen, Namin Wang, Lantian Li, Dong Wang 0013 |
ICASSP | 2 |
| 2024 | Fast Subsurface EM Simulation Based on the Mixed Finite-Element Method and Second-Order Arnoldi AlgorithmabstractA fast and accurate frequency sweep method combining the mixed finite element method (MFEM) with the second-order Arnoldi algorithm is proposed to characterize the EM material properties for high-loss geoelectromagnetic problems. The MFEM exploits the tree–cotree splitting technique for spatial discretization to eliminate the spurious modes arising from the null space of the curl operator. The modal order reduction method based on the second-order Arnoldi algorithm is employed to generate a set of orthogonal basis of the projection subspace for the quadratic eigenvalue problem (QEP) with both the conduction current and displacement current considered. The efficiency and accuracy of the method are verified by comparing with commercial software HFSS for lossy dielectric characterization in subsurface. Junhui Chen, Jie Liu 0051, Qing Huo Liu |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | Using full-scale feature fusion for self-supervised indoor depth estimation
Deqiang Cheng 0001, Junhui Chen, Chen Lv 0002, Chenggong Han |
Multim. Tools Appl. | 2 |
| 2024 | vEpiNet: A multimodal interictal epileptiform discharge detection method based on video and electroencephalogram dataabstractTo enhance deep learning-based automated interictal epileptiform discharge (IED) detection, this study proposes a multimodal method, vEpiNet, that leverages video and electroencephalogram (EEG) data. Datasets comprise 24 931 IED (from 484 patients) and 166 094 non-IED 4-second video-EEG segments. The video data is processed by the proposed patient detection method, with frame difference and Simple Keypoints (SKPS) capturing patients' movements. EEG data is processed with EfficientNetV2. The video and EEG features are fused via a multilayer perceptron. We developed a comparative model, termed nEpiNet, to test the effectiveness of the video feature in vEpiNet. The 10-fold cross-validation was used for testing. The 10-fold cross-validation showed high areas under the receiver operating characteristic curve (AUROC) in both models, with a slightly superior AUROC (0.9902) in vEpiNet compared to nEpiNet (0.9878). Moreover, to test the model performance in real-world scenarios, we set a prospective test dataset, containing 215 h of raw video-EEG data from 50 patients. The result shows that the vEpiNet achieves an area under the precision-recall curve (AUPRC) of 0.8623, surpassing nEpiNet's 0.8316. Incorporating video data raises precision from 70% (95% CI, 69.8%-70.2%) to 76.6% (95% CI, 74.9%-78.2%) at 80% sensitivity and reduces false positives by nearly a third, with vEpiNet processing one-hour video-EEG data in 5.7 min on average. Our findings indicate that video data can significantly improve the performance and precision of IED detection, especially in prospective real clinic testing. It suggests that vEpiNet is a clinically viable and effective tool for IED analysis in real-world applications. Weifang Gao, Junhui Chen, Zi Liang, Gonglin Yuan, Heyang Sun, Qing Li 0001, Liri Jin, Xiangqin Zhou, Chaoyue Dai, Haibo He, Yisu Dong, Liying Cui |
Neural Networks | 4 |
| 2023 | MICN: Multi-scale Local and Global Context Modeling for Long-term Series Forecasting
Jian Peng 0002, Feihu Huang 0002, Jince Wang, Junhui Chen, Yifei Xiao |
ICLR | 5 |
| 2023 | A Multi-Scale Attentive Transformer for Multi-Instrument Symbolic Music Generation
Xipin Wei, Junhui Chen, Zirui Zheng, Li Guo 0004, Lantian Li, Dong Wang 0013 |
INTERSPEECH | 2 |