Lei Li 0050

dblp:13/7007-50 · DBLP profile ↗
← Back
29ranked-venue papers
11as first author
29since 2021 · last 2026
0000-0002-2929-0828ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 4 first-author · 15 since 2021Artificial intelligence and machine learning · 13 · 2 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 4 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Theory of computation · 2 · 2 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multiple Human Motion Understanding
abstract
We introduce LLaMMo (Large Language and Multi-Person Motion Assistant), the first instruction-tuning multimodal framework tailored for multi-human motion analysis. LLaMMo incorporates a novel human-centric and social-temporal learner that models and fuses both intra-person dynamics and inter-person dependencies, yielding robust, context-aware representations of complex group behaviors while maintaining low computational overhead. To support LLaMMo, we construct LLaVerse, a large-scale dataset with fine-grained manual annotations covering diverse multi-person activities spanning daily social interaction and professional team sports. Built on top of LLaVerse, we also propose LLaMI-Bench, a dedicated benchmark for evaluating multi-human behavior understanding across motion and video modalities. Extensive experiments demonstrate that LLaMMo consistently outperforms baselines in understanding multi-person interactions under low-latency settings, with notable gains in both social and sport-specific contexts.
Lei Li 0050, Sen Jia 0003, Jenq-Neng Hwang
AAAI1
2026 3DSceneEditor: Controllable 3D Scene Editing with Gaussian Splatting
abstract
The creation of 3D scenes has traditionally been both labor-intensive and costly, requiring designers to meticulously configure 3D assets and environments. Recent advancements in generative AI, including text-to-3D and image-to-3D methods, have dramatically reduced the complexity and cost of this process. However, current techniques for editing complex 3D scenes continue to rely on generally interactive multi-step, 2D-to-3D projection methods and diffusion-based techniques, which often lack precision in control and hamper interactive-rate performance. In this work, we propose 3DSceneEditor, a fully 3D-based paradigm for interactive-rate, precise editing of intricate 3D scenes using Gaussian Splatting. Unlike conventional methods, 3DSceneEditor operates through a streamlined 3D pipeline, enabling direct Gaussian-based manipulation for efficient, high-quality edits based on input prompts. The proposed framework (i) integrates a pre-trained instance segmentation model for semantic labeling; (ii) employs a zero-shot grounding approach with CLIP to align target objects with user prompts; and (iii) applies scene modifications, such as object addition, repositioning, recoloring, replacing, and removal—directly on Gaussians. Extensive experimental results show that 3DSceneEditor surpasses existing state-of-the-art techniques in terms of both editing precision and efficiency, establishing a new benchmark for efficient and interactive 3D scene customization.
Ziyang Yan, Yihua Shao, Minwen Liao, Siyu Chen 0021, Nan Wang 0041, Muyuan Lin, Jenq-Neng Hwang, Hao Zhao 0002, Fabio Remondino, Lei Li 0050
WACV10
2026 Robust multi-domain digital pathology image segmentation via joint balancing representation learning
abstract
Multi-domain learning (MDL) seeks to mitigate performance degradation caused by domain-specific feature shifts between training and testing environments. In breast cancer digital pathology, such shifts result from variations in staining protocols, imaging devices, and tissue preparation. Existing MDL methods focus on aligning feature discrepancies across domains, often neglecting inter-domain feature balance caused by data distribution disparities. Addressing this balance is essential for effective cross-domain learning in breast cancer pathology, particularly in clinically diverse settings. Additionally, leveraging multi-source training data to enhance model adaptability across pathological domains remains challenging. We introduce a Joint Training Strategy (JTS) and a novel breast cancer digital pathology dataset with expert annotations to capture pathological heterogeneity across multiple domains. We propose Differential Ratio Integration with Twin-domain Training (DRIFT) for multi-source digital pathology image segmentation, addressing domain adaptation through: (1) PRISM-DD, a dynamic data distribution mechanism that balances domain contributions to optimize segmentation, and (2) AMBiCoL, a multi-level loss function integrating adaptive boundary masks with bidirectional consistency modeling to enhance generalization. Experiments on two breast cancer digital pathology datasets demonstrate that DRIFT outperforms state-of-the-art methods, supporting its potential for robust, scalable multi-domain segmentation in computational pathology. Code is available at https://github.com/Joycecc123/DRIFT .
Qiaoyi Xu, Afzan Adam, Azizi Abdullah, Tao Chen 0030, Adam J. Shephard, Patsy Ng Pei Sze, Noraidah Masir, Lei Li 0050, Reena Rahayu
Expert Syst. Appl.9
2026 MDS array codes with low disk I/O and small repair bandwidth
Lei Li 0050, Chenhao Ying 0001, Yuanyuan Dong 0002, Jie Li 0002, Yuan Luo 0003
Frontiers Comput. Sci.1
2026 Domain Adaptation for Speaker Verification Using Optimal Transport With Pseudo Label
abstract
Domain gap often degrades the performance of speaker verification (SV) systems when the statistical distributions of training data and real-world test speech are mismatched. Channel variation is a primary factor causing this gap, including bandwidth changes, background noise and encoding, etc. Although various domain adaptation algorithms could be applied to handle this domain gap problem, most algorithms could not take the complex distribution structure in domain alignment with discriminative learning. In this paper, we propose a novel unsupervised domain adaptation method for speaker verification, i.e., Joint Partial Optimal Transport with Pseudo Label (JPOT-PL), to alleviate the domain mismatch problem. Leveraging the geometric-aware distance metric of optimal transport in distribution alignment and speaker consistency in speech distribution, we further design a pseudo label-based discriminative learning where the pseudo label can be regarded as a new type of speaker label derived from the optimal coupling. With the JPOT-PL, we carry out experiments on the SV channel and lingual domain adaptation with VoxCeleb, LibriSpeech, CNCeleb, and AISHELL-2. Experiments show our method reduces EER by up to 30% compared with several state-of-the-art domain adaptation algorithms.
Jianguo Wei, Wenhuan Lu, Lei Li 0050, Xugang Lu
IEEE Trans. Inf. Forensics Secur.4
2025 Position-Aware Guided Point Cloud Completion with CLIP Model
abstract
Point cloud completion aims to recover partial geometric and topological shapes caused by equipment defects or limited viewpoints. Current methods either solely rely on the 3D coordinates of the point cloud to complete it or incorporate additional images with well-calibrated intrinsic parameters to guide the geometric estimation of the missing parts. Although these methods have achieved excellent performance by directly predicting the location of complete points, the extracted features lack fine-grained information regarding the location of the missing area. To address this issue, we propose a rapid and efficient method to expand an unimodal framework into a multimodal framework. This approach incorporates a position-aware module designed to enhance the spatial information of the missing parts through a weighted map learning mechanism. In addition, we establish a Point-Text-Image triplet corpus PCI-TI and MVP-TI based on the existing unimodal point cloud completion dataset and use the pre-trained vision-language model CLIP to provide richer detail information for 3D shapes, thereby enhancing performance. Extensive quantitative and qualitative experiments demonstrate that our method outperforms state-of-the-art point cloud completion methods.
Feng Zhou 0007, Ju Dai, Lei Li 0050, Junliang Xing
AAAI4
2025 The Role of Deductive and Inductive Reasoning in Large Language Models
abstract
Chengkun Cai, Xu Zhao, Haoliang Liu, Zhongyu Jiang, Tianfang Zhang, Zongkai Wu, Jenq-Neng Hwang, Lei Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Chengkun Cai, Haoliang Liu, Zhongyu Jiang, Tianfang Zhang, Zongkai Wu, Jenq-Neng Hwang, Lei Li 0050
ACL (1)8
2025 Human Motion Instruction Tuning
abstract
This paper presents LLaMo (Large Language and Human Motion Assistant), a multimodal framework for human motion instruction tuning. In contrast to conventional instruction-tuning approaches that convert non-linguistic inputs, such as video or motion sequences, into language tokens, LLaMo retains motion in its native form for instruction tuning. This method preserves motion-specific details that are often diminished in tokenization, thereby improving the model’s ability to interpret complex human behaviors. By processing both video and motion data alongside textual inputs, LLaMo enables a flexible, human-centric analysis. Experimental evaluations across high-complexity domains, including human behaviors and professional activities, indicate that LLaMo effectively captures domain-specific knowledge, enhancing comprehension and prediction in motion-intensive scenarios. We hope LLaMo offers a foundation for future multimodal AI systems with broad applications, from sports analytics to behavioral prediction.
Lei Li 0050, Sen Jia 0003, Zhongyu Jiang, Feng Zhou 0007, Ju Dai, Tianfang Zhang, Zongkai Wu, Jenq-Neng Hwang
CVPR1
2025 Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation
abstract
In 3D speech-driven facial animation generation, existing methods commonly employ pre-trained self-supervised audio models as encoders. However, due to the prevalence of phonetically similar syllables with distinct lip shapes in language, these near-homophone syllables tend to exhibit significant coupling in self-supervised audio feature spaces, leading to the averaging effect in subsequent lip motion generation. To address this issue, this paper proposes a plug-and-play semantic decorrelation module—Wav2Sem. This module extracts semantic features corresponding to the entire audio sequence, leveraging the added semantic information to decorrelate audio encodings within the feature space, thereby achieving more expressive audio features. Extensive experiments across multiple Speech-driven models indicate that the Wav2Sem module effectively decouples audio features, significantly alleviating the averaging effect of phonetically similar syllables in lip shape generation, thereby enhancing the precision and naturalness of facial animations. Our source code is available at https://github.com/wslh852/Wav2Sem.git.
Ju Dai, Xin Zhao 0025, Feng Zhou 0007, JunJun Pan, Lei Li 0050
CVPR6
2025 Learning an Efficient Optimizer via Hybrid-Policy Sub-Trajectory Balance
abstract
Recent advances in generative modeling enable neural networks to generate weights without relying on gradient-based optimization. However, current methods are limited by issues of over-coupling and long-horizon. The former tightly binds weight generation with task-specific objectives, thereby limiting the flexibility of the learned optimizer. The latter leads to inefficiency and low accuracy during inference, caused by the lack of local constraints. In this paper, we propose Lo-Hp, a decoupled two-stage weight generation framework that enhances flexibility through learning various optimization policies. It adopts a hybrid-policy sub-trajectory balance objective, which integrates on-policy and off-policy learning to capture local optimization policies. Theoretically, we demonstrate that learning solely local optimization policies can address the long-horizon issue while enhancing the generation of global optimal weights. In addition, we validate Lo-Hp’s superior accuracy and inference efficiency in tasks that require frequent weight updates, such as transfer learning, few-shot learning, domain generalization, and large language model adaptation.
Yunchuan Guan, Yu Liu 0040, Ke Zhou 0001, Sen Jia 0003, Zhiqi Shen 0001, Tao Chen 0030, Jenq-Neng Hwang, Lei Li 0050
ECAI11
2025 Is Meta-Learning Out? Rethinking Unsupervised Few-Shot Classification with Limited Entropy
abstract
Meta-learning is a powerful paradigm for tackling few-shot tasks. However, recent studies indicate that models trained with the whole-class training strategy can achieve comparable performance to those trained with meta-learning in few-shot classification tasks. To demonstrate the value of meta-learning, we establish an entropy-limited supervised setting for fair comparisons. Through both theoretical analysis and experimental validation, we establish that meta-learning has a tighter generalization bound compared to whole-class training. We unravel that meta-learning is more efficient with limited entropy and is more robust to label noise and heterogeneous tasks, making it well-suited for unsupervised tasks. Based on these insights, We propose MINO, a meta-learning framework designed to enhance unsupervised performance. MINO utilizes the adaptive clustering algorithm DBSCAN with a dynamic head for unsupervised task construction and a stability-based meta-scaler for robustness against label noise. Extensive experiments confirm its effectiveness in multiple unsupervised few-shot and zero-shot tasks.
Yunchuan Guan, Yu Liu 0040, Ke Zhou 0001, Zhiqi Shen 0001, Jenq-Neng Hwang, Serge J. Belongie, Lei Li 0050
ICCV7
2025 AU-Blendshape for Fine-Grained Stylized 3D Facial Expression Manipulation
abstract
While 3D facial animation has made impressive progress, challenges still exist in realizing fine-grained stylized 3D facial expression manipulation due to the lack of appropriate datasets. In this paper, we introduce the AUBlendSet, a 3D facial dataset based on AU-Blendshape representation for fine-grained facial expression manipulation across identities. AUBlendSet is a blendshape data collection based on 32 standard facial action units (AUs) across 500 identities, along with an additional set of facial postures annotated with detailed AUs. Based on AUBlendSet, we propose AUBlendNet to learn AU-Blendshape basis vectors for different character styles. AUBlendNet predicts, in parallel, the AU-Blendshape basis vectors of the corresponding style for a given identity mesh, thereby achieving stylized 3D emotional facial manipulation. We comprehensively validate the effectiveness of AUBlendSet and AUBlendNet through tasks such as stylized facial expression manipulation, speech-driven emotional facial animation, and emotion recognition data augmentation. Through a series of qualitative and quantitative experiments, we demonstrate the potential and importance of AUBlendSet and AUBlendNet in 3D facial animation tasks. To the best of our knowledge, AUBlendSet is the first dataset, and AUBlendNet is the first network for continuous 3D facial expression manipulation for any identity through facial AUs. Our source code is available at https://github.com/wslh852/AUBlendNet.git.
Ju Dai, Feng Zhou 0007, Kaida Ning, Lei Li 0050, JunJun Pan
ICCV5
2025 In-Context Meta LoRA Generation
abstract
Low-rank Adaptation (LoRA) has demonstrated remarkable capabilities for task specific fine-tuning. However, in scenarios that involve multiple tasks, training a separate LoRA model for each one results in considerable inefficiency in terms of storage and inference. Moreover, existing parameter generation methods fail to capture the correlations among these tasks, making multi-task LoRA parameter generation challenging. To address these limitations, we propose In-Context Meta LoRA (ICM-LoRA), a novel approach that efficiently achieves task-specific customization of large language models (LLMs). Specifically, we use training data from all tasks to train a tailored generator, Conditional Variational Autoencoder (CVAE). CVAE takes task descriptions as inputs and produces task-aware LoRA weights as outputs. These LoRA weights are then merged with LLMs to create task-specialized models without the need for additional fine-tuning. Furthermore, we utilize in-context meta-learning for knowledge enhancement and task mapping, to capture the relationship between tasks and parameter distributions. As a result, our method achieves more accurate LoRA parameter generation for diverse tasks using CVAE. ICM-LoRA enables more accurate LoRA parameter reconstruction than current parameter reconstruction methods and is useful for implementing task-specific enhancements of LoRA parameters. At the same time, our method occupies 283MB, only 1% storage compared with the original LoRA. The code is available at https://github.com/YihuaJerry/ICM-LoRA.
Yihua Shao, Minxi Yan, Yang Liu 0360, Siyu Chen 0021, Xinwei Long, Ziyang Yan, Lei Li 0050, Nicu Sebe, Hao Tang 0005, Yan Wang 0068, Hao Zhao 0002, Mengzhu Wang, Jingcai Guo
IJCAI8
2025 Constructions of Binary Cooperative MSR Codes with Optimal Access Bandwidth
abstract
Minimum storage regenerating (MSR) codes are extensively studied in the literature to reduce the network bandwidth consumed during node repair. In this paper, we focus on repairing multiple node failures and construct binary cooperative MSR codes with optimal access bandwidth. Specifically, we present explicit constructions of the codes by designing the parity-check matrices over a special polynomial ring ${\mathcal{R}} = {{\mathbb{F}}_2}[x]/\left({1 + x + \cdots + {x^{p - 1}}}\right)$ where p is a prime. The obtained codes with length n and dimension k achieve the lower bound on repair bandwidth under cooperative repair model with optimal access property for any 2 ≤ h ≤ r and k+1 ≤ d ≤ n – h where h and d are the number of failure nodes and helper nodes, respectively. Moreover, the computation operations involved in encoding, decoding and nodes repair are only exclusive ORs and cyclic shifts, resulting in less CPU overhead compared with the complex multiplication operations over finite fields.
Lei Li 0050, Xinchun Yu, Yaqian Zhang 0002, Yuan Luo 0003
ITW1
2025 MoCount: Motion-Based Repetitive Action Counting
abstract
Existing action counting methods typically rely on pixel-based changes within videos, leading to high computational redundancy and low accuracy due to the limited spatial sensitivity. To address these challenges, we introduce MoCount, the first framework that leverages 3D motion representations for counting tasks. MoCount significantly reduces computational overhead and improves counting accuracy, benefiting from the simplicity of motion representation and strong spatial sensitivity. Specifically, we utilize a motion estimator to convert video subjects into 3D motion data. A motion encoder, combined with a Sparse Spatial-Temporal module, is then applied to extract robust human body representations, yielding precise counting results. Extensive experiments on the RepCount and UCFRep datasets show that MoCount achieves state-of-the-art performance, reducing inference latency by approximately 2-3 times compared to existing video counting models. These advantages position MoCount as a leading solution for real-world action counting applications.
Ruocheng Gu, Sen Jia 0003, Yule Ma, Jinqin Zhong, Jenq-Neng Hwang, Lei Li 0050
ACM Multimedia6
2025 Graph Canvas for Controllable 3D Scene Generation
abstract
Spatial intelligence is fundamental to AI systems that interact with the physical world, particularly in 3D scene generation and spatial comprehension. Current layout generation in 3D scene synthesis remains highly complex, often constrained by predefined datasets and limited dynamic adaptation to changing spatial relationships. In this paper, we propose GraphCanvas3D, a flexible, query-driven framework for controllable 3D scene generation. Unlike traditional methods that require retraining and predefined input masks for modifications, GraphCanvas3D provides a training-free solution supporting the generation of diverse scenes-both indoor and outdoor-through free manipulation of objects and scene elements. Our framework employs hierarchical, graph-driven scene descriptions, representing spatial elements as graph nodes and establishing coherent relationships among objects in 3D environments. The decoupled object representation enables flexible, on-the-fly scene adjustments and dynamic, customizable scene creation. Experimental results and user studies demonstrate that GraphCanvas3D improves usability, adaptability, and generalization across various 3D scene generation tasks, offering a powerful tool for scalable and diverse scene synthesis.
Sen Jia 0003, Jingzhe Shi, Can Jin, Zongkai Wu, Jenq-Neng Hwang, Lei Li 0050
ACM Multimedia8
2024 Construction of Binary Cooperative MSR Codes with Multiple Repair Degrees
Lei Li 0050, Xinchun Yu, Yaqian Zhang 0002, Yuanyuan Dong 0002, Chenhao Ying 0001, Yuan Luo 0003
COCOON (2)1
2024 Constructions of Binary MDS Array Codes with Optimal Cooperative Repair Bandwidth
abstract
Erasure codes are widely implemented in distributed storage systems to provide high fault tolerance with small storage overhead. Maximum distance separable codes are an common choice as they achieve the optimal tradeoff between fault tolerance and storage overhead. In this paper, we focus on the repair of multiple erasures of binary MDS array codes. Specifically, we present constructions of binary MDS array codes with optimal cooperative repair bandwidth by stacking multiple Blaum-Roth code instances whose “evaluation points” are judiciously designed. The constructed array codes with length$n$and dimension$k$can achieve the optimal cooperative repair bandwidth for$2\leq h\leq n-k$and$k+1\leq d\leq n-h$where$h$and$d$are the numbers of failed nodes and helper nodes, respectively. As the codes are constructed on a special polynomial ring over binary field, computation operations involved in nodes repair and file reconstruction for these codes are only XORs and cyclic shifts. Moreover, due to the inherent parallel structure of the codes, both the encoding and decoding procedures can be finished in parallel, speeding up the computing process.
Lei Li 0050, Xinchun Yu, Chenhao Ying 0001, Yuanyuan Dong 0002, Yuan Luo 0003
ISIT1
2024 Back to Optimization: Diffusion-based Zero-Shot 3D Human Pose Estimation
abstract
Learning-based methods have dominated the 3D human pose estimation (HPE) tasks with significantly better performance in most benchmarks than traditional optimization-based methods. Nonetheless, 3D HPE in the wild is still the biggest challenge for learning-based models, whether with 2D-3D lifting, image-to-3D, or diffusion-based methods, since the trained networks implicitly learn camera intrinsic parameters and domain-based 3D human pose distributions and estimate poses by statistical average. On the other hand, the optimization-based methods estimate results case-by-case, which can predict more diverse and sophisticated human poses in the wild. By combining the advantages of optimization-based and learning-based methods, we propose the Zero-shot Diffusion-based Optimization (ZeDO) pipeline for 3D HPE to solve the problem of cross-domain and in-the-wild 3D HPE. Our multi-hypothesis ZeDO achieves state-of-the-art (SOTA) performance on Human3.6M, with minMPJPE 51.4mm, without training with any 2D-3D or image-3D pairs. Moreover, our single-hypothesis ZeDO achieves SOTA performance on 3DPW dataset with PA-MPJPE 40.3mm on cross-dataset evaluation, which even outperforms learning-based methods trained on 3DPW. Our code is available here: https://github.com/ipl-uw/ZeDO-Release.
Zhongyu Jiang, Zhuoran Zhou, Lei Li 0050, Wenhao Chai, Cheng-Yen Yang, Jenq-Neng Hwang
WACV3
2024 CPSeg: Finer-grained Image Semantic Segmentation via Chain-of-Thought Language Prompting
abstract
Natural scene analysis and remote sensing imagery offer immense potential for advancements in large-scale language-guided context-aware data utilization. This potential is particularly significant for enhancing performance in downstream tasks such as object detection and segmentation with designed language prompting. In light of this, we introduce the CPSeg (Chain-of-Thought Language Prompting for Finer-grained Semantic Segmentation), an innovative framework designed to augment image segmentation performance by integrating a novel "Chain-of-Thought" process that harnesses textual information associated with images. This groundbreaking approach has been applied to a flood disaster scenario. CPSeg encodes prompt texts derived from various sentences to formulate a coherent chain-of-thought. We use a new vision-language dataset, FloodPrompt, which includes images, semantic masks, and corresponding text information. This not only strengthens the semantic understanding of the scenario but also aids in the key task of semantic segmentation through an interplay of pixel and text matching maps. Our qualitative and quantitative analyses validate the effectiveness of CPSeg.
Lei Li 0050
WACV1
2024 RPCANet: Deep Unfolding RPCA Based Infrared Small Target Detection
abstract
Deep learning (DL) networks have achieved remarkable performance in infrared small target detection (ISTD). However, these structures exhibit a deficiency in interpretability and are widely regarded as black boxes, as they disregard domain knowledge in ISTD. To alleviate this issue, this work proposes an interpretable deep network for detecting infrared dim targets, dubbed RPCANet. Specifically, our approach formulates the ISTD task as sparse target extraction, low-rank background estimation, and image reconstruction in a relaxed Robust Principle Component Analysis (RPCA) model. By unfolding the iterative optimization updating steps into a deep-learning framework, time-consuming and complex matrix calculations are replaced by theory-guided neural networks. RPCANet detects targets with clear interpretability and preserves the intrinsic image feature, instead of directly transforming the detection task into a matrix decomposition problem. Extensive experiments substantiate the effectiveness of our deep unfolding framework and demonstrate its trustworthy results, surpassing baseline methods in both qualitative and quantitative evaluations. Our source code is available at https://github.com/fengyiwu98/RPCANet.
Fengyi Wu, Tianfang Zhang, Lei Li 0050, Yian Huang, Zhenming Peng
WACV3
2024 MDS array codes with efficient repair and small sub-packetization level
Lei Li 0050, Xinchun Yu, Chenhao Ying 0001, Yuanyuan Dong 0002, Yuan Luo 0003
Des. Codes Cryptogr.1
2024 Constructions of Binary MDS Array Codes With Optimal Repair/Access Bandwidth
abstract
Maximum distance separable (MDS) codes are commonly deployed in distributed storage systems as they provide the maximum failure tolerance for some given redundancy. The repair problem of MDS codes has drawn much attention and various constructions of MDS array codes with optimal repair bandwidth have been proposed in the last decade. However, few of the existing codes are constructed over the binary field. In this paper, we propose new constructions of binary MDS array codes with optimal repair (or access) bandwidth for single-node failure. Specifically, by stacking multiple Blaum-Roth code instances of which the parity-check matrices are judiciously designed, we obtain three families of binary MDS array codes with optimal repair bandwidth; using the permutation matrices as building blocks, we also construct two families of binary MDS array codes with optimal access bandwidth. Moreover, error-resilient capability while achieving the lower bound on repair (or access) bandwidth is obtained when the number of helper nodes d < n - 1. All the codes in this paper are constructed over a particular ring of binary polynomials. Consequently, computation operations involved in the encoding, decoding and node repair procedures for these codes are only XORs and cyclic shifts, avoiding complex multiplications and divisions over large finite fields.
Lei Li 0050, Xinchun Yu, Yuanyuan Dong 0002, Yuan Luo 0003
IEEE Trans. Commun.1
2024 Incentive Mechanism for Uncertain Tasks Under Differential Privacy
abstract
Mobile crowd sensing (MCS) has emerged as an increasingly popular sensing paradigm due to its cost-effectiveness. This approach relies on platforms to outsource tasks to participating workers when prompted by task publishers. Although incentive mechanisms have been devised to foster widespread participation in MCS, most of them focus only on static tasks (i.e., tasks for which the timing and type are known in advance) and do not protect the privacy of worker bids. In a dynamic and resource-constrained environment, tasks are often uncertain (i.e., the platform lacks a priori knowledge about the tasks) and worker bids may be vulnerable to inference attacks. This paper presents an incentive mechanism HERALD*, that takes into account the uncertainty and hidden bids of tasks without real-time constraints. Theoretical analysis reveals that HERALD* satisfies a range of critical criteria, including truthfulness, individual rationality, differential privacy, low computational complexity, and low social cost. These properties are then corroborated through a series of evaluations.
Xikun Jiang, Chenhao Ying 0001, Lei Li 0050, Boris Düdder, Haiqin Wu, Haiming Jin, Yuan Luo 0003
IEEE Trans. Serv. Comput.3
2023 New Constructions of Binary MDS Array Codes with Optimal Repair Bandwidth
abstract
Maximum distance separable codes are commonly used in large-scale distributed storage systems since they achieve the maximum fault tolerance for some given redundancy. In this paper, we focus on the repair problem of MDS codes. By stacking multiple Blaum-Roth code instances of which the parity-check matrices are judiciously designed, we construct three families of binary MDS array codes with optimal repair bandwidth for single-node failure. Specifically, codes in the first family achieve the lower bound on repair bandwidth where the number of helper nodes d is n − 1; The second family of codes are optimal-repair for any fixed d and the third are for multiple values of d simultaneously. Moreover, the last two also possess error-resilient capability and can achieve the corresponding optimal repair bandwidth. All the codes in this paper are constructed on a particular polynomial ring over binary field. Consequently, computation operations involved in node repair and file reconstruction for these codes are only XORs and cyclic shifts, avoiding complex multiplications and divisions over large finite fields.
Lei Li 0050, Chenhao Ying 0001, Yuanyuan Dong 0002, Yuan Luo 0003
ISIT1
2023 Mask-FPAN: Semi-supervised face parsing in the wild with de-occlusion and UV GAN
Lei Li 0050, Tianfang Zhang, Zhongfeng Kang, Xikun Jiang
Comput. Graph.1
2023 Optimization-inspired Cumulative Transmission Network for image compressive sensing
Tianfang Zhang, Lei Li 0050, Zhenming Peng
Knowl. Based Syst.2
2023 Multi-scale pseudo labeling for unsupervised deep edge detection
abstract
Deep learning currently rules edge detection. However, the impressive progress heavily relies on high-quality manually annotated labels which require a significant amount of labor and time. In this study, we propose a novel unsupervised learning framework for deep edge detection. It adopts a gradient-based method to generate scale-dependent pseudo edge maps, which match with the hierarchical structure of deep networks. It leverages both the representation learning capability of deep learning, and the simplicity of traditional methods. Experiments on three popular data sets show that the proposed method can suppress non-object edges and reduce the gap with its supervised counterpart due to the introduction of information of various scales and smoothing strategy.
Changsheng Zhou, Hongxin Wang, Lei Li 0050, Stefan Oehmcke, Junmin Liu
Knowl. Based Syst.4
2022 Deep learning based 3D point cloud regression for estimating forest biomass
abstract
Knowledge of forest biomass stocks and their development is important for implementing effective climate change mitigation measures. Remote sensing using airborne LiDAR can be used to measure vegetation structure at large scale. We present deep learning systems for predicting wood volume, above-ground biomass (AGB), and subsequently above-ground carbon stocks directly from airborne LiDAR point clouds. Specifically, we devise different neural network architectures for point cloud regression and evaluate them on remote sensing data of areas for which AGB estimates have been obtained from field measurements in a national forest inventory. Our adaptation of Minkowski convolutional neural networks for regression gave the best results. The deep neural networks produced significantly more accurate wood volume, AGB, and carbon estimates compared to state-of-the-art approaches operating on basic statistics of the point clouds. In contrast to other methods, no digital terrain model is required. We expect this finding to have a strong impact on LiDAR-based analyses of terrestrial ecosystem dynamics.
Stefan Oehmcke, Lei Li 0050, Jaime C. Revenga, Thomas Nord-Larsen, Katerina Trepekli, Fabian Gieseke, Christian Igel
SIGSPATIAL/GIS2