VLDB 2026 Research / reviewers in the wild / expert
Yuheng Lu
dblp:155/3107
· DBLP profile ↗
18ranked-venue papers
5as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evaluating the Expressive Appropriateness of Speech in Rich ContextsabstractTianrui Wang, Ziyang Ma, Yizhou Peng, Haoyu Wang, Zhikang Niu, Zikang Huang, Yihao Wu, Yi-Wen Chao, Yu Jiang, Yuheng Lu, Guanrou Yang, Xuanchen Li, Hexin Liu, Chunyu Qiang, Cheng Gong, Yifan Yang, Tianchi Liu, Junyu Wang, Nana Hou, Meng Ge, Fuming You, Yang Wei, Zhongqian Sun, Hu Haifeng, Xiaobao Wang, Eng Siong Chng, Xie Chen, Longbiao Wang, Jianwu Dang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Tianrui Wang, Ziyang Ma 0001, Yizhou Peng, Zhikang Niu, Zikang Huang, Yi-Wen Chao, Yuheng Lu, Guanrou Yang, Xuanchen Li, Hexin Liu, Chunyu Qiang, Yifan Yang 0005, Tianchi Liu 0004, Nana Hou, Meng Ge, Fuming You, Zhongqian Sun, Haifeng Hu 0009, Xiaobao Wang, Chng Eng Siong, Xie Chen 0001, Longbiao Wang, Jianwu Dang 0001 |
ACL (1) | 10 |
| 2026 | Partial-Correlation Learning for Large Language Models with Skip-Tuning
Yuheng Lu, Zuhe Song, Caixia Yuan |
ICPR (15) | 1 |
| 2026 | A novel spatial downscaling algorithm based on deep learning considering geographical spatial heterogeneity and nonlinear changes: a case study of the Yangtze River Basin
Chuanjiang Luo, Lilu Cui, Jing Xiang, Yuheng Lu, Haoyang Guo, Jiachun An |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | Neuron Counting for Macaque Mesoscopic Brain Connectivity ResearchabstractPrecise quantification and localization of tracer-labeled neurons are essential for unraveling brain connectivity patterns and constructing a mesoscopic brain connectome atlas in macaques. However, methodological challenges and limitations in dataset development have impeded this scientific progress. This work introduced the Macaque Fluorescently Labeled Neurons (MFN) dataset, derived from retrograde tracing on three rhesus macaques. The dataset, meticulously annotated by six specialists, includes 1,600 images and 33,411 high-quality neuron annotations. Leveraging this dataset, we developed a Dense Convolutional Attention U-Net (DAUNet) cell counting model. By integrating Dense Convolutional blocks and a multi-scale attention module, the model exhibits robust feature extraction and representation capabilities while maintaining low complexity. On the MFN dataset, DAUNet achieved a Mean Absolute Error of 0.97 for cell counting and an F1-score of 96.29% for cell localization, outperforming several benchmark models. Extensive validation across four additional public datasets demonstrated the robust generalization ability of the model. Furthermore, the trained model was applied to quantify labeled neurons of a macaque brain, mapping the input connectivity patterns of two adjacent subregions in the lateral prefrontal cortex. This work provides a training dataset and algorithmic resource that advances mesoscopic brain connectivity research in macaques. The MFN dataset and source code are available at https://github.com/Gendwar/DAUnet. Zhenwei Dong, Weiyang Shi, Yuheng Lu, Xiaoxiao Hou, Hongji Sun, Zhengyi Yang 0002, Tianzi Jiang |
IEEE Trans. Medical Imaging | 4 |
| 2025 | Controlled Low-Rank Adaptation with Subspace Regularization for Continued Training on Large Language ModelsabstractLarge language models (LLMs) exhibit remarkable capabilities in natural language processing but face catastrophic forgetting when learning new tasks, where adaptation to a new domain leads to a substantial decline in performance on previous tasks. In this paper, we propose Controlled LoRA (CLoRA), a subspace regularization method on LoRA structure. Aiming to reduce the scale of output change while introducing minimal constraint on model capacity, CLoRA imposes constraints on the direction of updating matrix’s null space. Experimental results on one-stage LLM finetuning tasks and continual learning settings highlight the superiority of CLoRA as an effective parameter-efficient finetuning method with catastrophic forgetting mitigating. Further investigation for model parameters indicates that CLoRA effectively balances the trade-off between model capacity and degree of forgetting. The code for implementing CLoRA will be publicly available. Yuheng Lu, Bingshuo Qian, Caixia Yuan, Huixing Jiang |
ACL (1) | 1 |
| 2025 | Discrete Unit-based Low-latency Multi-lingual Speech Synthesis for LIMMITS'25 ChallengeabstractIn this paper, we present the system developed by our team, CCATTS, for the LIMMITS’25 challenge, focusing on few-shot and zero-shot TTS. We adopt a two-stage TTS strategy. In track 1, we fine-tune the pre-trained ZMM-TTS model and successfully achieve multilingual low-latency TTS. In track 2, we propose a token-based framework by modifying the first-stage model of ZMM-TTS, disentangling speech features into four types of discrete tokens—Content, Acoustic, Emotion, and Speaker—and integrating it with our designed token2wav module. This module consists of a HiFi-GAN-style decoder, an acoustic refiner, and U-Net flow matching, to generate high-quality speech. The official competition results demonstrate that our method achieves strong performance in both tracks. Tianrui Wang, Chunyu Qiang, Qiuyu Liu, Yuheng Lu, Xiaobao Wang, Longbiao Wang, Jianwu Dang 0001 |
ICASSP | 7 |
| 2024 | TS-AI: A deep learning pipeline for multimodal subject-specific parcellation with task contrasts synthesis
Chengyi Li, Yuheng Lu, Yue Cui 0005 |
Medical Image Anal. | 2 |
| 2024 | Multimodal Connectivity-Based Individual Parcellation and Analysis for Humans and Rhesus MonkeysabstractIndividual brains vary greatly in morphology, connectivity and organization. Individualized brain parcellation is capable of precisely localizing subject-specific functional regions. However, most individualization approaches have examined single modalities of data and have not generalized to nonhuman primates. The present study proposed a novel multimodal connectivity-based individual parcellation (MCIP) method, which optimizes within-region homogeneity, spatial continuity and similarity to a reference atlas with the fusion of personal functional and anatomical connectivity. Comprehensive evaluation demonstrated that MCIP outperformed state-of-the-art multimodal individualization methods in terms of functional and anatomical homogeneity, predictability of cognitive measures, heritability, reproducibility and generalizability across species. Comparative investigation showed a higher topographic variability in humans than that in macaques. Therefore, MCIP provides improved accurate and reliable mapping of brain functional regions over existing methods at an individual level across species, and could facilitate comparative and translational neuroscience research. Yue Cui 0005, Chengyi Li, Yuheng Lu, Luqi Cheng, Long Cao, Tianzi Jiang |
IEEE Trans. Medical Imaging | 3 |
| 2024 | BAI-Net: Individualized Anatomical Cerebral Cartography Using Graph Neural NetworkabstractBrain atlas is an important tool in the diagnosis and treatment of neurological disorders. However, due to large variations in the organizational principles of individual brains, many challenges remain in clinical applications. Brain atlas individualization network (BAI-Net) is an algorithm that subdivides individual cerebral cortex into segregated areas using brain morphology and connectomes. The presented method integrates group priors derived from a population atlas, adjusts areal probabilities using the context of connectivity fingerprints derived from the fiber-tract embedding of tractography, and provides reliable and explainable individualized brain areas across multiple sessions and scanners. We demonstrate that BAI-Net outperforms the conventional iterative clustering approach by capturing significantly heritable topographic variations in individualized cartographies. The topographic variability of BAI-Net cartographies has shown strong associations with individual variability in brain morphology, connectivity as well as higher relationship on individual cognitive behaviors and genetics. This study provides an explainable framework for individualized brain cartography that may be useful in the precise localization of neuromodulation and treatments on individual brains. Yu Zhang 0115, Hantian Zhang, Luqi Cheng, Zhengyi Yang 0002, Yuheng Lu, Weiyang Shi, Wen Li 0021, Junjie Zhuo, Jiaojian Wang, Lingzhong Fan, Tianzi Jiang |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | PiMAE: Point Cloud and Image Interactive Masked Autoencoders for 3D Object DetectionabstractMasked Autoencoders learn strong visual representations and achieve state-of-the-art results in several independent modalities, yet very few works have addressed their capabilities in multi-modality settings. In this work, we focus on point cloud and RGB image data, two modalities that are often presented together in the real world, and explore their meaningful interactions. To improve upon the cross-modal synergy in existing works, we propose Pi-MAE, a self-supervised pre-training framework that promotes 3D and 2D interaction through three aspects. Specifically, we first notice the importance of masking strategies between the two sources and utilize a projection module to complementarily align the mask and visible tokens of the two modalities. Then, we utilize a well-crafted two-branch MAE pipeline with a novel shared decoder to promote cross-modality interaction in the mask tokens. Finally, we design a unique cross-modal reconstruction module to enhance representation learning for both modalities. Through extensive experiments performed on large-scale RGB-D scene understanding benchmarks (SUN RGB-D and ScannetV2), we discover it is nontrivial to interactively learn point-image features, where we greatly improve multiple 3D detectors, 2D detectors, and few-shot classifiers by 2.9%, 6.7%, and 2.4%, respectively. Code is available at https://github.com/BLVLab/PiMAE. Anthony Chen, Renrui Zhang, Zihan Wang 0011, Yuheng Lu, Yandong Guo, Shanghang Zhang |
CVPR | 5 |
| 2023 | Open-Vocabulary Point-Cloud Object Detection without 3D AnnotationabstractThe goal of open-vocabulary detection is to identify novel objects based on arbitrary textual descriptions. In this paper, we address open-vocabulary 3D point-cloud detection by a dividing-and-conquering strategy, which involves: 1) developing a point-cloud detector that can learn a general representation for localizing various objects, and 2) connecting textual and point-cloud representations to enable the detector to classify novel object categories based on text prompting. Specifically, we resort to rich image pretrained models, by which the point-cloud detector learns localizing objects under the supervision of predicted 2D bounding boxes from 2D pretrained detectors. Moreover, we propose a novel de-biased triplet cross-modal contrastive learning to connect the modalities of image, point-cloud and text, thereby enabling the point-cloud detector to benefit from vision-language pretrained models, i.e., CLIP. The novel use of image and vision-language pretrained models for point-cloud detectors allows for open-vocabulary 3D object detection without the need for 3D annotations. Experiments demonstrate that the proposed method improves at least 3.03 points and 7.47 points over a wide range of baselines on the ScanNet and SUN RGB-D datasets, respectively. Furthermore, we provide a comprehensive analysis to explain why our approach works. Code is available at https://github.com/lyhdet/OV-3DET Yuheng Lu, Chenfeng Xu, Xiaobao Wei, Masayoshi Tomizuka, Kurt Keutzer, Shanghang Zhang |
CVPR | 1 |
| 2023 | Can Language Really Understand Depth?
Fangping Chen, Yuheng Lu |
ICONIP (13) | 2 |
| 2022 | Enhancing and Dissecting Crowd Counting by Synthetic DataabstractIn this article, we propose a simulated crowd counting dataset CrowdX, which has a large scale, accurate labeling, parameterized realization, and high fidelity. The experimental results of using this dataset as data enhancement show that the performance of the proposed streamlined and efficient benchmark network ESA-Net can be improved by 8.4%. The other two classic heterogeneous architectures MCNN and CSRNet pre-trained on CrowdX also show significant performance improvements. Considering many influencing factors determine performance, such as background, camera angle, human density, and resolution. Although these factors are important, there is still a lack of research on how they affect crowd counting. Thanks to the CrowdX dataset with rich annotation information, we conduct a large number of data-driven comparative experiments to analyze these factors. Our research provides a reference for a deeper understanding of the crowd counting problem and puts forward some useful suggestions in the actual deployment of the algorithm. Chengyang Li 0001, Yuheng Lu, Yuan Li 0014, Huizhu Jia |
ICASSP | 3 |
| 2022 | Directed Mix Contrast for Lidar Point Cloud SegmentationabstractComprehensive and real-time scene understanding are crucial for autonomous driving, where LiDAR semantic segmentation plays an indispensable role. However, existing segmentation algorithms are exposed with extremely imbalanced dataset. Besides, due to the sparseness of point cloud, some hard negative samples are intrinsically similar. In this paper, we propose a contrastive learning framework named Directed Mix Contrast (DMC) for LiDAR point cloud segmentation. There are two key contributions in this framework. Firstly, a simple pre-processing step is proposed to maintain the balance between classes. In addition, to deal with the hard negative samples, contrastive learning strategy is introduced to facilitate feature learning. More specifically, we propose a directed mix-up training strategy to take control of the contrasting procedure. To demonstrate the effectiveness of DMC, we conduct experiments on two large-scale LiDAR Segmentation datasets and achieve mIoU of 71.0% on SemanticK-ITTI, which outperforms most existing methods. Furthermore, we extend our method to several backbone networks, and the results show good generalization ability of DMC. Yuheng Lu, Fangping Chen, Ziwei Zhang 0003, Fan Yang 0053 |
ICME | 1 |
| 2021 | Real-time active detection of targets and path planning using UAVsabstractThis article proposes a new method that enables Unmanned Aerial Vehicles (UAVs) to actively find targets and shoot photographs of them in an unknown environment, while successfully avoiding surrounding obstacles and planning optimize routes. Owing to the limited computing ability on the UAVs, we obtained the point cloud data of surrounding objects, and selected the best segmentation method of the point cloud to perform real-time semantic segmentation on the collected point cloud data. The point cloud data with semantic attributes were merged into voxels. We reconstruct the real-time distance and angle between the surface of obstacles and the surrounding obstacles through Euclidean Signed Distance Fields (ESDFs), and adjust the gimbal angle and focal length of UAVs and use the two-dimensional image recognition to shoot the photographs of the target precisely. Considering the increasing scale of UAVs power inspections, we can improve the efficiency of fine inspections of power transmission lines by using the method we proposed. Fangping Chen, Yuheng Lu, Yunyi Li |
ICRA | 2 |
| 2019 | Cross-X Learning for Fine-Grained Visual CategorizationabstractRecognizing objects from subcategories with very subtle differences remains a challenging task due to the large intra-class and small inter-class variation. Recent work tackles this problem in a weakly-supervised manner: object parts are first detected and the corresponding part-specific features are extracted for fine-grained classification. However, these methods typically treat the part-specific features of each image in isolation while neglecting their relationships between different images. In this paper, we propose Cross-X learning, a simple yet effective approach that exploits the relationships between different images and between different network layers for robust multi-scale feature learning. Our approach involves two novel components: (i) a cross-category cross-semantic regularizer that guides the extracted features to represent semantic parts and, (ii) a cross-layer regularizer that improves the robustness of multi-scale features by matching the prediction distribution across multiple layers. Our approach can be easily trained end-to-end and is scalable to large datasets like NABirds. We empirically analyze the contributions of different components of our approach and demonstrate its robustness, effectiveness and state-of-the-art performance on five benchmark datasets. Code is available at \url{https://github.com/cswluo/CrossX}. Wei Luo 0006, Xitong Yang, Xianjie Mo, Yuheng Lu, Larry Davis 0001, Jun Li 0027, Jian Yang 0003, Ser-Nam Lim |
ICCV | 4 |
| 2016 | Learning to Predict miRNA-mRNA Interactions from AGO CLIP Sequencing and CLASH DataabstractRecent technologies like AGO CLIP sequencing and CLASH enable direct transcriptome-wide identification of AGO binding and miRNA target sites, but the most widely used miRNA target prediction algorithms do not exploit these data. Here we use discriminative learning on AGO CLIP and CLASH interactions to train a novel miRNA target prediction model. Our method combines two SVM classifiers, one to predict miRNA-mRNA duplexes and a second to learn a binding model of AGO's local UTR sequence preferences and positional bias in 3'UTR isoforms. The duplex SVM model enables the prediction of non-canonical target sites and more accurately resolves miRNA interactions from AGO CLIP data than previous methods. The binding model is trained using a multi-task strategy to learn context-specific and common AGO sequence preferences. The duplex and common AGO binding models together outperform existing miRNA target prediction algorithms on held-out binding data. Open source code is available at https://bitbucket.org/leslielab/chimiric. Yuheng Lu, Christina S. Leslie |
PLoS Comput. Biol. | 1 |
| 2014 | Anonymous Camera for Privacy ProtectionabstractPrivacy protection in the surveillance video data has received great attention. Although tremendous works have been proposed to provide effective privacy protection techniques, most of the algorithms are based on post-processing that deletes, obscures or encrypts the privacy information after privacy-included raw data are recorded. Consequently, they are vulnerable to raw data leak out, which may lead to unauthorized use. Therefore, it is imperative to develop a new privacy protection scheme which is capable of excluding any privacy information at the video recording phase. In this paper, we propose an anonymous camera aiming to protect the privacy of individuals at the video capturing phase by optical masking technique. It effectively reduces the risk of raw data leakage because no privacy information will be recorded by the camera. We implemented a prototype camera, which consists of an infrared camera, a RGB camera and a liquid crystal on silicon (LCoS) device. We introduce optical design and performance of the anonymous camera, the masking algorithm as well as the calibration methodology. Experimental results demonstrate that our prototype anonymous camera can perform accurate real time masking of the face for privacy protection. Yuheng Lu, Hajime Nagahara, Rin-Ichiro Taniguchi |
ICPR | 2 |