Shuo Lu

dblp:49/53 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Security and privacy · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 RAGAR: Retrieval Augmented Personalized Image Generation Guided by Recommendation
abstract
Personalized image generation is crucial for improving the user experience, as it renders reference images into preferred ones according to user visual preferences. Although effective, existing methods face two main issues. First, existing methods treat all items in the user's historical sequence equally when extracting user preferences, overlooking the varying semantic similarities between historical items and the reference item. Disproportionately high weights for low-similarity items distort user visual preferences for the reference item. Second, existing methods heavily rely on consistency between generated and reference images to optimize generation, which leads to underfitting user preferences and hinders personalization. To address these issues, we propose Retrieval Augmented Personalized Image GenerAtion guided by Recommendation (RAGAR). Our approach uses a retrieval mechanism to assign different weights to historical items according to their similarities to the reference item, thereby extracting more refined users' visual preferences for the reference item. Then we introduce a novel rank task based on the multi-modal ranking model to optimize the personalization of the generated images instead of forcing depend on consistency. Extensive experiments and human evaluations on three real-world datasets demonstrate that RAGAR achieves significant improvements in both personalization and semantic metrics compared to five baselines.
Run Ling, Wenji Wang, Yuting Liu 0003, Guibing Guo, Quanwei Zhang, Yexing Xu, Shuo Lu, Yihua Shao, Linying Jiang, Xingwei Wang 0001
AAAI9
2026 GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging
abstract
Beyond scratch coding, exploiting large-scale code repositories (e.g., GitHub) for practical tasks is vital in real-world software development, yet current benchmarks rarely evaluate code agents in such authentic, workflow-driven scenarios. To bridge this gap, we introduce GitTaskBench, a benchmark designed to systematically assess this capability via 54 realistic tasks across 7 modalities and 7 domains. Each task pairs a relevant repository with an automated, human-curated evaluation harness specifying practical success criteria. Beyond measuring execution and task success, we also propose the alpha-value metric to quantify the economic benefit of agent performance, which integrates task success rates, token cost, and average developer salaries. Experiments across three state-of-the-art agent frameworks with multiple advanced LLMs show that leveraging code repositories for complex task solving remains challenging: even the best-performing system, OpenHands+Claude 3.7, solves only 48.15% of tasks. Error analysis attributes over half of failures to seemingly mundane yet critical steps like environment setup and dependency resolution, highlighting the need for more robust workflow management and increased timeout preparedness. By releasing GitTaskBench, we aim to drive progress and attention toward repository-aware code reasoning, execution, and deployment---moving agents closer to solving complex, end-to-end real-world tasks.
Ziyi Ni, Huacan Wang, Shuo Lu, Wang You, Zhenheng Tang, Sen Hu 0005, Bo Li 0117, Binxing Jiao, Daxin Jiang, Yuntao Du 0001
AAAI4
2026 Beyond Boundaries: Leveraging Vision Foundation Models for Source-Free Object Detection
abstract
Source-Free Object Detection (SFOD) aims to adapt a source-pretrained object detector to a target domain without access to source data. However, existing SFOD methods predominantly rely on internal knowledge from the source model, which limits their capacity to generalize across domains and often results in biased pseudo-labels, thereby hindering both transferability and discriminability. In contrast, Vision Foundation Models (VFMs), pretrained on massive and diverse data, exhibit strong perception capabilities and broad generalization, yet their potential remains largely untapped in the SFOD setting. In this paper, we propose a novel SFOD framework that leverages VFMs as external knowledge sources to jointly enhance feature alignment and label quality. Specifically, we design three VFM-based modules: (1) Patch-weighted Global Feature Alignment (PGFA) distills global features from VFMs using patch-similarity–based weighting to enhance global feature transferability; (2) Prototype-based Instance Feature Alignment (PIFA) performs instance-level contrastive learning guided by momentum-updated VFM prototypes; and (3) Dual-source Enhanced Pseudo-label Fusion (DEPF) fuses predictions from detection VFMs and teacher models via an entropy-aware strategy to yield more reliable supervision. Extensive experiments on six benchmarks demonstrate that our method achieves state-of-the-art SFOD performance, validating the effectiveness of integrating VFMs to simultaneously improve transferability and discriminability.
Huizai Yao, Sicheng Zhao, Pengteng Li, Shuo Lu, Weiyu Guo, Yunfan Lu, Yijie Xu, Hui Xiong 0001
AAAI5
2026 What If Consensus Lies? Selective-Complementary Reinforcement Learning at Test Time
abstract
Test-Time Reinforcement Learning (TTRL) enables Large Language Models (LLMs) to enhance reasoning capabilities on unlabeled test streams by deriving pseudo-rewards from majority voting consensus.However, existing TTRL methods rely exclusively on positive pseudo-labeling strategies.Such reliance becomes vulnerable under challenging scenarios where answer distributions are highly dispersed, resulting in weak consensus that inadvertently reinforces incorrect trajectories as supervision signals.In this paper, we propose SCRL (Selective-Complementary Reinforcement Learning), a robust test-time reinforcement learning framework that effectively mitigates label noise amplification.SCRL develops Selective Positive Pseudo-Labeling, which enforces strict consensus criteria to filter unreliable majorities.Complementarily, SCRL introduces Entropy-Gated Negative Pseudo-Labeling, the first negative supervision mechanism in TTRL, to reliably prune incorrect trajectories based on generation uncertainty.Extensive experiments on multiple reasoning benchmarks demonstrate that SCRL achieves substantial improvements over baselines, while maintaining robust generalization and training stability under constrained rollout budgets.Our code is available at https://github.com/Jasper
Jian Liang 0001, Yanbo Wang 0004, Shuo Lu, Ran He 0001, Tieniu Tan
ACL (1)4
2025 Uni-Layout: Integrating Human Feedback in Unified Layout Generation and Evaluation
Shuo Lu, Yanyin Chen, Fengheng Li, Jingjing Lv, Junjie Shen 0008, Ching Law, Jian Liang 0001
ACM Multimedia1
2025 RepoMaster: Autonomous Exploration and Understanding of GitHub Repositories for Complex Task Solving
abstract
The ultimate goal of code agents is to solve complex tasks autonomously. Although large language models (LLMs) have made substantial progress in code generation, real-world tasks typically demand full-fledged code repositories rather than simple scripts. Building such repositories from scratch remains a major challenge. Fortunately, GitHub hosts a vast, evolving collection of open-source repositories, which developers frequently reuse as modular components for complex tasks. Yet, existing frameworks like OpenHands and SWE-Agent still struggle to effectively leverage these valuable resources. Relying solely on README files provides insufficient guidance, and deeper exploration reveals two core obstacles: overwhelming information and tangled dependencies of repositories, both constrained by the limited context windows of current LLMs. To tackle these issues, we propose RepoMaster, an autonomous agent framework designed to explore and reuse GitHub repositories for solving complex tasks. For efficient understanding, RepoMaster constructs function-call graphs, module-dependency graphs, and hierarchical code trees to identify essential components, providing only identified core elements to the LLMs rather than the entire repository. During autonomous execution, it progressively explores related components using our exploration tools and prunes information to optimize context usage. Evaluated on the adjusted MLE-bench, RepoMaster achieves a 110\% relative boost in valid submissions over the strongest baseline OpenHands. On our newly released GitTaskBench, RepoMaster lifts the task-pass rate from 40.7% to 62.9% while reducing token usage by 95%. Our code and demonstration materials are publicly available at https://github.com/QuantaAlpha/RepoMaster.
Huacan Wang, Ziyi Ni, Shuo Lu, Sen Hu 0005, Jiaye Lin, Yifu Guo, Yuntao Du 0001
NeurIPS4
2025 Frustratingly Easy Feature Reconstruction for Out-of-Distribution Detection
Yingsheng Wang, Shuo Lu, Jian Liang 0001, Aihua Zheng, Ran He 0001
PRCV (9)2
2025 Adaptive hyperparameter optimization for author name disambiguation
abstract
Abstract In the process of author name disambiguation (AND), varying characteristics and noise of different blocks significantly impact disambiguation performance. In this paper, we propose a block‐based adaptive hyperparameter optimization method that assigns optimal hyperparameters to each block without altering the original AND model structure. Based on this, a random forest model is trained using the optimized results to fit the relationship between the block's data features and its optimal hyperparameters, thereby enabling the prediction of hyperparameters for new blocks. Empirical studies on 6 state‐of‐the‐art AND algorithms, 11 public datasets, and a manually labeled dataset of China's information and communication technology (ICT) industry patents demonstrate that the proposed method significantly outperforms the original algorithms across multiple standard performance evaluation metrics (Cluster F1/Pairwise F1, B‐Cubed F1, and K metrics). The results of the random forest regression indicate that the selected 16 features effectively predict the optimal hyperparameters. Further analysis reveals a power‐law relationship between relative block size and both relative performance and relative optimized performance across all datasets and evaluation metrics, and the relative performance improvement of the adaptive hyperparameter optimization algorithm is particularly significant for smaller blocks. These findings provide theoretical support and practical guidance for the development of AND algorithms.
Shuo Lu
J. Assoc. Inf. Sci. Technol.1
2025 Source-Free Object Detection With Detection Transformer
abstract
Source-Free Object Detection (SFOD) enables knowledge transfer from a source domain to an unsupervised target domain for object detection without access to source data. Most existing SFOD approaches are either confined to conventional object detection (OD) models like Faster R-CNN or designed as general solutions without tailored adaptations for novel OD architectures, especially Detection Transformer (DETR). In this paper, we introduce Feature Reweighting ANd Contrastive Learning NetworK (FRANCK), a novel SFOD framework specifically designed to perform query-centric feature enhancement for DETRs. FRANCK comprises four key components: 1) an Objectness Score-based Sample Reweighting (OSSR) module that computes attention-based objectness scores on multi-scale encoder feature maps, reweighting the detection loss to emphasize less-recognized regions; 2) a Contrastive Learning with Matching-based Memory Bank (CMMB) module that integrates multi-level features into memory banks, enhancing class-wise contrastive learning; 3) an Uncertainty-weighted Query-fused Feature Distillation (UQFD) module that improves feature distillation through prediction quality reweighting and query feature fusion; and 4) an improved self-training pipeline with a Dynamic Teacher Updating Interval (DTUI) that optimizes pseudo-label quality. By leveraging these components, FRANCK effectively adapts a source-pre-trained DETR model to a target domain with enhanced robustness and generalization. Extensive experiments on several widely used benchmarks demonstrate that our method achieves state-of-the-art performance, highlighting its effectiveness and compatibility with DETR-based SFOD models.
Huizai Yao, Sicheng Zhao, Shuo Lu, Hui Chen 0013, Tengfei Xing, Chenggang Yan 0001, Jianhua Tao 0001, Guiguang Ding
IEEE Trans. Image Process.3
2025 Spatiotemporal Microstate Dynamics of Spike-Free Scalp EEG Offer a Potential Biomarker for Refractory Temporal Lobe Epilepsy
abstract
Refractory temporal lobe epilepsy (TLE) is one of the most frequently observed subtypes of epilepsy and endangers more than 50 million people world-wide. Although electroencephalogram (EEG) had been widely recognized as a classic tool to screen and diagnose epilepsy, for many years it heavily relied on identifying epileptic discharges and epileptogenic zone localization, which however, limits the understanding of refractory epilepsy due to the network nature of this disease. This work hypothesizes that the microstate dynamics based on resting-state scalp EEG can offer an additional network depiction of the disease and provide potential complementary evaluation tool for the TLE even without detectable epileptic discharges on EEG. We propose a novel framework for EEG microstate spatial-temporal dynamics (EEG-MiSTD) analysis based on machine learning to comprehensively model millisecond-changing whole-brain network dynamics. With only 100 seconds of resting-state EEG even without epileptic discharges, this approach successfully distinguishes TLE patients from healthy controls and is related to the lateralization of epileptic focus. Besides, microstate temporal and spatial features are found to be widely related to clinical parameters, which further demonstrate that TLE is a network disease. A preliminary exploration suggests that the spatial topography is sensitive to the following surgical outcomes. From such a new perspective, our results suggest that spatiotemporal microstate dynamics is potentially a biomarker of the disease. The developed EEG-MiSTD framework can probably be considered as a general tool to examine dynamical brain network disruption in a user-friendly way for other types of epilepsy.
Zelin Chen, Ruiyan Feng, N. U. Farrukh Hameed, Liang Chen 0023, Shuo Lu
IEEE Trans. Medical Imaging10
2023 Android Malicious Application Detection Based on Improved Mayfly Algorithm
abstract
With the rapid development of the Internet and mobile terminals, Android is one of the most popular mobile operating systems. However, the proliferation of Android malicious applications is also very serious, so it is necessary to detect and deal with Android applications in advance, and how to make effective selection among many features is a crucial process in malicious application detection. In this paper, an Android malicious application detection model is established based on the improved mayfly algorithm with reference to related Android malicious detection methods for Android platform applications. By effectively selecting the features, the optimal combination of features is obtained to optimize the classification results, so as to improve the detection performance of Android malicious application detection.The static analysis is used to extract the features of Android applications, and the malicious application detection model is examined by various classification algorithms, and the results confirm the feasibility and superiority of the proposed Android malicious application detection method based on the improved mayfly algorithm.
Yinzhen Wei, Shuo Lu
TrustCom2
2022 Exploring the relationship between children's facial emotion processing characteristics and speech communication ability using deep learning on eye tracking and speech performance measures
Zelin Chen, Guoxin Qiu, Zhuanggui Chen, Leyan Gao, Shuo Lu
Comput. Speech Lang.9
2008 Securing Telehealth Applications in a Web-Based e-Health Portal
abstract
Telehealth applications can deliver medical services to patients at remote locations using telecommunications technologies, such as the Internet. At the same time, such applications also pose unique security challenges. First, the trust issue becomes more severe due to the lack of visual proofs in telehealth applications. The public key infrastructure (PKI) is insufficient for providing the same kind of trust a patient may attain during a face-to-face service. Second, telehealth services, such as tele-monitoring or tele-consultant, naturally demand a systematic organization of users, roles, resources, and flows of information. Existing access control mechanisms in an e-health system are usually incapable of dealing with such workflow-based services. This paper provides cost-efficient solutions to those issues in the context of a Web-based e-health portal system. First, we propose a PKI-like infrastructure for establishing trust between users using biometrics-based authentication and hierarchies of trust. Second, we develop an access control method for workflow-based telehealth services using a rule-based module already available in the portal system.
Shuo Lu, Yuan Hong 0001, Lingyu Wang 0001, Rachida Dssouli
ARES2
2008 Preserving Privacy in E-health Systems Using Hippocratic Databases
abstract
Safeguarding patientspsila private information is one of the most challenging issues in the design and implementation of modern e-Health systems. Recent advances in Hippocratic Databases (HDB) show a promising direction towards the enforcement of privacy policies in e-Health systems. This paper tackles issues in applying the HDB design to e-Health systems. More specifically, we design an architecture for integrating APPEL preferences with HDB; we extend the original HDB design to support fine-grained privacy authorizations demanded by patients; we adapt the design to a multi-dimensional model; we also propose a design for hierarchical authorizations. Finally, we discuss implementation issues and justify our designs with experimental results.
Yuan Hong 0001, Shuo Lu, Lingyu Wang 0001, Rachida Dssouli
COMPSAC2