EDBT 2026 Demo / reviewers in the wild / expert
Xinfeng Li
dblp:04/8388
· DBLP profile ↗
61ranked-venue papers
13as first author
43since 2021 · last 2026
0000-0002-9686-4369ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 17 since 2021Security and privacy · 16 · 6 first-author · 16 since 2021Computer networks · 13 · 7 first-author · 2 since 2021Systems, architecture and hardware · 5 · 1 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dynamic Generation of Multi LLM Agents Communication Topologies with Graph Diffusion ModelsabstractEric Hanchen Jiang, Levina Li, Frank Wan, Xiao Liang, Sophia Yin, Yuchen Wu, Xinfeng Li, Yizhou Sun, Wei Wang, Kai-Wei Chang, Ying Nian Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Eric Hanchen Jiang, Levina Li, Frank Wan, Sophia Yin, Xinfeng Li, Yizhou Sun, Kai-Wei Chang 0001, Ying Nian Wu |
ACL (1) | 7 |
| 2026 | Mitigating Over-Refusal in Aligned Large Language Models via Inference-Time Activation EnergyabstractEric Hanchen Jiang, Weixuan Ou, Run Liu, Shengyuan Pang, Guancheng Wan, Ranjie Duan, Wei Dong, Kai-Wei Chang, XiaoFeng Wang, Ying Nian Wu, Xinfeng Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Eric Hanchen Jiang, Weixuan Ou, Run Liu 0005, Shengyuan Pang, Guancheng Wan, Ranjie Duan, Wei Dong 0007, Kai-Wei Chang 0001, Xiaofeng Wang 0001, Ying Nian Wu, Xinfeng Li |
ACL (1) | 11 |
| 2026 | Buster: Implanting Semantic Backdoor Into Text Encoder to Mitigate NSFW Content Generation
Xiaojun Chen 0004, Yuexin Xuan, Zhendong Zhao, Xinfeng Li, Xiaojun Jia, Xiaofeng Wang 0001 |
DASFAA (5) | 5 |
| 2026 | EmoRAG: Evaluating RAG Robustness to Symbolic PerturbationsabstractRetrieval-Augmented Generation (RAG) systems are increasingly central to robust AI, enhancing large language model (LLM) faithfulness by incorporating external knowledge. However, our study unveils a critical, overlooked vulnerability: their profound susceptibility to subtle symbolic perturbations, particularly through near-imperceptible emotional icons (e.g., "(@_@)") that can catastrophically mislead retrieval, termed EmoRAG. We demonstrate that injecting a single emoticon into a query makes it nearly 100% likely to retrieve semantically unrelated texts, which contain a matching emoticon. Our extensive experiment across general question-answering and code domains, using a range of state-of-the-art retrievers and generators, reveals three key findings: (I) Single-Emoticon Disaster: Minimal emoticon injections cause maximal disruptions, with a single emoticon almost 100% dominating RAG output. (II) Positional Sensitivity: Placing an emoticon at the beginning of a query can cause severe perturbation, with F1-Scores exceeding 0.92 across all datasets. (III) Parameter-Scale Vulnerability: Counterintuitively, models with larger parameters exhibit greater vulnerability to the interference. We provide an in-depth analysis to uncover the underlying mechanisms of these phenomena. Furthermore, we raise a critical concern regarding the robustness assumption of current RAG systems, envisioning a threat scenario where an adversary exploits this vulnerability to manipulate the RAG system. We evaluate standard defenses and find them insufficient against EmoRAG. To address this, we propose targeted defenses, analyzing their strengths and limitations in mitigating emoticon-based perturbations. Finally, we outline future directions for building robust RAG systems. Xinyun Zhou, Xinfeng Li, Yinan Peng, Ming Xu 0006, Xuanwang Zhang, Yidong Wang 0003, Xiaojun Jia, Kun Wang 0056, Qingsong Wen, XiaoFeng Wang 0001, Wei Dong 0007 |
KDD (1) | 2 |
| 2026 | WebCloak: Characterizing and Mitigating Threats From LLM-Driven Web Agents as Intelligent Scrapers
Xinfeng Li, Tianze Qiu, Yingbin Jin, Lixu Wang, Hanqing Guo, Xiaojun Jia, Xiaofeng Wang 0001, Wei Dong 0007 |
SP | 1 |
| 2026 | The Person Behind the Sound: Demystifying Audio Private Attribute Profiling Via Multimodal Large Language Models
Lixu Wang, Kaixiang Yao, Xinfeng Li, Haoyao Li, Xiaofeng Wang 0001, Wei Dong 0007 |
SP | 3 |
| 2026 | ENCHTABLE: Unified Safety Alignment Transfer in Fine-Tuned Large Language ModelsabstractMany machine learning models are fine-tuned from large language models (LLMs) to achieve high performance in specialized domains like code generation, biomedical analysis, and mathematical problem solving. However, this fine-tuning process often introduces a critical vulnerability: the systematic degradation of safety alignment, undermining ethical guidelines and increasing the risk of harmful outputs. Addressing this challenge, we introduce EnchTable, a novel framework designed to transfer and maintain safety alignment in downstream LLMs without requiring extensive retraining. EnchTable leverages a Neural Tangent Kernel (NTK)-based safety vector distillation method to decouple safety constraints from task-specific reasoning, ensuring compatibility across diverse model architectures and sizes. Additionally, our interference-aware merging technique effectively balances safety and utility, minimizing performance compromises across various task domains. We implemented a fully functional prototype of EnchTable on three different task domains and three distinct LLM architectures, and evaluated its performance through extensive experiments on eleven diverse datasets, assessing both utility and model safety. Our evaluations include LLMs from different vendors, demonstrating EnchTable's generalization capability. Furthermore, EnchTable exhibits robust resistance to static and dynamic jailbreaking attacks, outperforming vendor-released safety models in mitigating adversarial prompts. Comparative analyses with six parameter modification methods and two inference-time alignment baselines reveal that EnchTable achieves a significantly lower unsafe rate, higher utility score, and universal applicability across different task domains. Additionally, we validate EnchTable can be seamlessly integrated into various deployment pipelines without significant overhead. Jialin Wu 0001, Kecen Li, Xinfeng Li, XiaoFeng Wang 0001, Cheng Hong 0001 |
SP | 4 |
| 2026 | Multi-view clustering with tensor-aligned dynamic anchor graphs
Xinyue Deng, Xinfeng Li, Yuying Zhu 0014, Zhenwen Ren |
Neurocomputing | 3 |
| 2026 | Critical Information Only: A Content Privacy-Preserving Framework for Detecting Audio DeepfakesabstractText-to-Speech (TTS) and Voice Conversion (VC) models have exhibited remarkable performance in generating realistic and natural audio. However, their dark side, audio deepfake poses a significant threat to both society and individuals. Existing countermeasures largely focus on determining the genuineness of speech based on complete original audio recordings, which however often contain private content. This oversight may refrain deepfake detection from many applications, particularly in scenarios involving sensitive information like business secrets. In this paper, we propose SafeEar, a novel framework that aims to detect deepfake audios without relying on accessing the speech content within. Our key idea is to devise a neural audio codec into a novel decoupling model that well separates the semantic and acoustic information from audio samples, and only use the acoustic information (e.g., prosody and timbre) for deepfake detection. In this way, no semantic content will be exposed to the detector. To overcome the challenge of identifying diverse deepfake audio without semantic clues, we enhance our deepfake detector with real-world augmentation, such as codecs and reverbs. Extensive experiments conducted on five benchmark datasets demonstrate SafeEar's effectiveness in detecting various deepfake techniques with an equal error rate (EER) down to 2.41%. Simultaneously, it shields f ive-language speech content from being deciphered by both machine and human auditory analysis, demonstrated by word error rates (WERs) all above 93.74% and our user study. Furthermore, our benchmark constructed for anti-deepfake and anti-content recovery evaluation helps provide a basis for future research in the realms of audio privacy preservation and deepfake detection. Xinfeng Li, Yifan Zheng 0001, Chen Yan 0001, Kai Li 0047, Chang Zeng, Xiaoyu Ji 0001, Wenyuan Xu 0001 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2026 | PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image ModelsabstractRecent text-to-image (T2I) models have exhibited remarkable performance in generating high-quality images from text descriptions. However, these models are vulnerable to misuse, particularly generating not-safe-for-work (NSFW) content, such as sexually explicit, violent, political, and disturbing images, raising serious ethical concerns. In this work, we present PromptGuard, a novel content moderation technique that draws inspiration from the system prompt mechanism in large language models (LLMs) for safety alignment. Unlike LLMs, T2I models lack a direct interface for enforcing behavioral guidelines. Our key idea is to optimize a safety soft prompt that functions as an implicit system prompt within the T2I model’s textual embedding space. This universal soft prompt (P∗) directly moderates NSFW inputs, enabling safe yet realistic image generation without altering the inference efficiency or requiring proxy models.We further enhance its reliability and helpfulness through a divide-and-conquer strategy, which optimizes category-specific soft prompts and combines them into holistic safety guidance. Extensive experiments across five datasets demonstrate that PromptGuard effectively mitigates NSFW content generation while preserving high-quality benign outputs. PromptGuard achieves 3.8 times faster than prior content moderation methods, surpassing eight state-of-the-art defenses. Rigorous evaluation using both multi-head classifiers and VLM-based guardrails confirms its robustness, achieving an optimal average unsafe ratios down to 5.84% and 6.18%, respectively. Our code and dataset are available at https://t2ipromptguard. github.io/. Lingzhi Yuan, Xinfeng Li, Chejian Xu, Guanhong Tao 0001, Xiaojun Jia, Yihao Huang 0001, Wei Dong 0007, Yang Liu 0003, Bo Li 0026 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Pierce the Mists, Greet the Sky: Decipher Knowledge Overshadowing via Knowledge Circuit AnalysisabstractLarge Language Models (LLMs), despite their remarkable capabilities, are hampered by hallucinations.A particularly challenging variant, knowledge overshadowing, occurs when one piece of activated knowledge inadvertently masks another relevant piece, leading to erroneous outputs even with high-quality training data.Current understanding of overshadowing is largely confined to inference-time observations, lacking deep insights into its origins and internal mechanisms during model training.Therefore, we introduce PHANTOMCIRCUIT, a novel framework designed to comprehensively analyze and detect knowledge overshadowing.By innovatively employing knowledge circuit analysis, PHANTOMCIRCUIT dissects the function of key components in the circuit and how the attention pattern dynamics contribute to the overshadowing phenomenon and its evolution throughout the training process.Extensive experiments demonstrate PHANTOMCIRCUIT 's effectiveness in identifying such instances, offering novel insights into this elusive hallucination and providing the research community with a new methodological lens for its potential mitigation.Our code can be found in https://github.com/halfmorepiece/PhantomCircuit. Haoming Huang, Jiahao Huo, Xin Zou 0001, Xinfeng Li, Kun Wang 0056, Xuming Hu |
EMNLP | 5 |
| 2025 | DynamicNER: A Dynamic, Multilingual, and Fine-Grained Dataset for LLM-based Named Entity RecognitionabstractHanjun Luo, Yingbin Jin, Yiran Wang, Xinfeng Li, Tong Shang, Xuecheng Liu, Ruizhe Chen, Kun Wang, Hanan Salam, Qingsong Wen, Zuozhu Liu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Hanjun Luo, Yingbin Jin, Xinfeng Li, Tong Shang, Xuecheng Liu, Ruizhe Chen, Kun Wang 0056, Hanan Salam, Qingsong Wen, Zuozhu Liu |
EMNLP | 4 |
| 2025 | Heuristic-Induced Multimodal Risk Distribution Jailbreak Attack for Multimodal Large Language Models
Xiaojun Jia, Ranjie Duan, Xinfeng Li, Yihao Huang 0001, Xiaoshuang Jia, Zhixuan Chu, Wenqi Ren |
ICCV | 4 |
| 2025 | A Vision for Access Control in LLM Agent Systems
Hongyi Cai, Xinfeng Li, Yijia Xu |
ICECCS | 4 |
| 2025 | Empowering Embodied Agents with Semantic Intelligence
Wenbing Tang 0001, Meilin Zhu, Fenghua Wu, Xinfeng Li, Yang Liu 0003 |
ICECCS | 4 |
| 2025 | Agent Behavior: The Regulatory Object of the Agent-Centric Online Ecosystem in Digital Age
Qiang Zhang 0057, Pei Yan, Yijia Xu, Xinfeng Li, Hongyi Cai, Chuanpo Fu, Yong Fang 0002, Yang Liu 0003 |
ICECCS | 4 |
| 2025 | A Survey on Trustworthy LLM Agents: Threats and CountermeasuresabstractWith the rapid evolution of Large Language Models (LLMs), LLMbased agents and Multi-agent Systems (MAS) have significantly expanded the capabilities of LLM ecosystems.This evolution stems from empowering LLMs with additional modules such as memory, tools, environment, and even other agents.However, this advancement has also introduced more complex issues of trustworthiness, which previous research focusing solely on LLMs could not cover.In this survey, we propose the TrustAgent framework, a comprehensive study on the trustworthiness of agents, characterized by modular taxonomy, multi-dimensional connotations, and * Miao Yu and Fanci Meng contribute equally to this paper. Fanci Meng, Xinyun Zhou, Shilong Wang 0002, Junyuan Mao, Linsey Pang, Tianlong Chen 0001, Kun Wang 0056, Xinfeng Li, Yongfeng Zhang 0003, Bo An 0001, Qingsong Wen |
KDD (2) | 9 |
| 2025 | The Eye of Sherlock Holmes: Uncovering User Private Attribute Profiling via Vision-Language Model Agentic Framework
Feiran Liu, Xinyi Huang 0015, Yinan Peng, Xinfeng Li, Lixu Wang, Yutong Shen, Ranjie Duan, Simeng Qin, Xiaojun Jia, Qingsong Wen, Wei Dong 0007 |
ACM Multimedia | 5 |
| 2025 | RACONTEUR: A Knowledgeable, Insightful, and Portable LLM-Powered Shell Command Explainer
Jiangyi Deng, Xinfeng Li, Yanjiao Chen, Yijie Bai, Haiqin Weng, Yan Liu 0069, Tao Wei 0002, Wenyuan Xu 0001 |
NDSS | 2 |
| 2025 | LightAntenna: Characterizing the Limits of Fluorescent Lamp-Induced Electromagnetic Interference
Fengchen Yang, Wenze Cui, Xinfeng Li, Chen Yan 0001, Xiaoyu Ji 0001, Wenyuan Xu 0001 |
NDSS | 3 |
| 2025 | Adversarial Attacks against Closed-Source MLLMs via Feature Optimal AlignmentabstractMultimodal large language models (MLLMs) remain vulnerable to transferable adversarial examples. While existing methods typically achieve targeted attacks by aligning global features—such as CLIP’s [CLS] token—between adversarial and target samples, they often overlook the rich local information encoded in patch tokens. This leads to suboptimal alignment and limited transferability, particularly for closed-source models. To address this limitation, we propose a targeted transferable adversarial attack method based on feature optimal alignment, called FOA-Attack, to improve adversarial transfer capability. Specifically, at the global level, we introduce a global feature loss based on cosine similarity to align the coarse-grained features of adversarial samples with those of target samples. At the local level, given the rich local representations within Transformers, we leverage clustering techniques to extract compact local patterns to alleviate redundant local features. We then formulate local feature alignment between adversarial and target samples as an optimal transport (OT) problem and propose a local clustering optimal transport loss to refine fine-grained feature alignment. Additionally, we propose a dynamic ensemble model weighting strategy to adaptively balance the influence of multiple models during adversarial example generation, thereby further improving transferability. Extensive experiments across various models demonstrate the superiority of the proposed method, outperforming state-of-the-art methods, especially in transferring to closed-source MLLMs. Xiaojun Jia, Sensen Gao, Simeng Qin, Tianyu Pang, Yihao Huang 0001, Xinfeng Li, Yiming Li 0004, Bo Li 0026, Yang Liu 0003 |
NeurIPS | 7 |
| 2025 | GuardReasoner-VL: Safeguarding VLMs via Reinforced ReasoningabstractTo enhance the safety of VLMs, this paper introduces a novel reasoning-based VLM guard model dubbed GuardReasoner-VL. The core idea is to incentivize the guard model to deliberatively reason before making moderation decisions via online RL.
First, we construct GuardReasoner-VLTrain, a reasoning corpus with 123K samples and 631K reasoning steps, spanning text, image, and text-image inputs.
Then, based on it, we cold-start our model's reasoning ability via SFT.
In addition, we further enhance reasoning regarding moderation through online RL.
Concretely, to enhance diversity and difficulty of samples, we conduct rejection sampling followed by data augmentation via the proposed safety-aware data concatenation.
Besides, we use a dynamic clipping parameter to encourage exploration in early stages and exploitation in later stages.
To balance performance and token efficiency, we design a length-aware safety reward that integrates accuracy, format, and token cost.
Extensive experiments demonstrate the superiority of our model.
Remarkably, it surpasses the runner-up by 19.27% F1 score on average, as shown in Figure 1.
We release data, code, and models (3B/7B) of GuardReasoner-VL: https://github.com/yueliu1999/GuardReasoner-VL. Yue Liu 0008, Shengfang Zhai, Mingzhe Du, Tri Cao, Hongcheng Gao, Xinfeng Li, Kun Wang 0056, Junfeng Fang, Jiaheng Zhang, Bryan Hooi |
NeurIPS | 8 |
| 2025 | AgentAuditor: Human-level Safety and Security Evaluation for LLM AgentsabstractDespite the rapid advancement of LLM-based agents, the reliable evaluation of their safety and security remains a significant challenge. Existing rule-based or LLM-based evaluators often miss dangers in agents' step-by-step actions, overlook subtle meanings, fail to see how small issues compound, and get confused by unclear safety or security rules. To overcome this evaluation crisis, we introduce AgentAuditor, a universal, training-free, memory-augmented reasoning framework that empowers LLM evaluators to emulate human expert evaluators. AgentAuditor constructs an experiential memory by having an LLM adaptively extract structured semantic features (e.g., scenario, risk, behavior) and generate associated chain-of-thought reasoning traces for past interactions. A multi-stage, context-aware retrieval-augmented generation process then dynamically retrieves the most relevant reasoning experiences to guide the LLM evaluator's assessment of new cases. Moreover, we developed ASSEBench, the first benchmark designed to check how well LLM-based evaluators can spot both safety risks and security threats. ASSEBench comprises 2293 meticulously annotated interaction records, covering 15 risk types across 29 application scenarios. A key feature of ASSEBench is its nuanced approach to ambiguous risk situations, employing "Strict" and "Lenient" judgment standards. Experiments demonstrate that AgentAuditor not only consistently improves the evaluation performance of LLMs across all benchmarks but also sets a new state-of-the-art in LLM-as-a-judge for agent safety and security, achieving human-level accuracy. Our work is openly accessible at https://github.com/Astarojth/AgentAuditor-ASSEBench. Hanjun Luo, Shenyu Dai, Chiming Ni, Xinfeng Li, Guibin Zhang, Kun Wang 0056, Tongliang Liu, Hanan Salam |
NeurIPS | 4 |
| 2025 | MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their MixabstractWe introduce MMAR, a new benchmark designed to evaluate the deep reasoning capabilities of Audio-Language Models (ALMs) across massive multi-disciplinary tasks. MMAR comprises 1,000 meticulously curated audio-question-answer triplets, collected from real-world internet videos and refined through iterative error corrections and quality checks to ensure high quality. Unlike existing benchmarks that are limited to specific domains of sound, music, or speech, MMAR extends them to a broad spectrum of real-world audio scenarios, including mixed-modality combinations of sound, music, and speech. Each question in MMAR is hierarchically categorized across four reasoning layers: Signal, Perception, Semantic, and Cultural, with additional sub-categories within each layer to reflect task diversity and complexity. To further foster research in this area, we annotate every question with a Chain-of-Thought (CoT) rationale to promote future advancements in audio reasoning. Each item in the benchmark demands multi-step deep reasoning beyond surface-level understanding. Moreover, a part of the questions requires graduate-level perceptual and domain-specific knowledge, elevating the benchmark's difficulty and depth. We evaluate MMAR using a broad set of models, including Large Audio-Language Models (LALMs), Large Audio Reasoning Models (LARMs), Omni Language Models (OLMs), Large Language Models (LLMs), and Large Reasoning Models (LRMs), with audio caption inputs. The performance of these models on MMAR highlights the benchmark's challenging nature, and our analysis further reveals critical limitations of understanding and reasoning capabilities among current models. These findings underscore the urgent need for greater research attention in audio-language reasoning, including both data and algorithm innovation. We hope MMAR will serve as a catalyst for future advances in this important but little-explored area. Ziyang Ma 0001, Yinghao Ma, Yanqiao Zhu 0003, Yi-Wen Chao, Yuanzhe Chen, Zhuo Chen 0006, Jian Cong, Keliang Li, Siyou Li, Xinfeng Li, Xiquan Li, Zheng Lian 0004, Yuzhe Liang, Minghao Liu 0003, Zhikang Niu, Tianrui Wang, Yuping Wang 0005, Yuxuan Wang 0002, Guanrou Yang, Jianwei Yu 0001, Ruibin Yuan, Zhisheng Zheng, Ziya Zhou, Haina Zhu, Wei Xue 0002, Emmanouil Benetos, Kai Yu 0004, Chng Eng Siong, Xie Chen 0001 |
NeurIPS | 14 |
| 2025 | MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video ScenariosabstractMultimodal Large Language Models (MLLMs) have achieved considerable accuracy in Optical Character Recognition (OCR) from static images. However, their efficacy in video OCR is significantly diminished due to factors such as motion blur, temporal variations, and visual effects inherent in video content. To provide clearer guidance for training practical MLLMs, we introduce MME-VideoOCR benchmark, which encompasses a comprehensive range of video OCR application scenarios. MME-VideoOCR features 10 task categories comprising 25 individual tasks and spans 44 diverse scenarios. These tasks extend beyond text recognition to incorporate deeper comprehension and reasoning of textual content within videos. The benchmark consists of 1,464 videos with varying resolutions, aspect ratios, and durations, along with 2,000 meticulously curated, manually annotated question-answer pairs. We evaluate 18 state-of-the-art MLLMs on MME-VideoOCR, revealing that even the best-performing model (Gemini-2.5 Pro) achieves only an accuracy of 73.7%. Fine-grained analysis indicates that while existing MLLMs demonstrate strong performance on tasks where relevant texts are contained within a single or few frames, they exhibit limited capability in effectively handling tasks that demand holistic video comprehension. These limitations are especially evident in scenarios that require spatio-temporal reasoning, cross-frame information integration, or resistance to language prior bias. Our findings also highlight the importance of high-resolution visual input and sufficient temporal coverage for reliable OCR in dynamic video scenarios. Yang Shi 0009, Huanqian Wang, Wulin Xie, Huanyao Zhang, Lijie Zhao, Yifan Zhang 0004, Xinfeng Li, Chaoyou Fu, Zhuoer Wen, Zhuoran Zhang 0003, Xinlong Chen, Bohan Zeng, Yushuo Guan, Zhang Zhang 0001, Liang Wang 0001, Haoxuan Li 0001, Zhouchen Lin, Yuanxing Zhang, Pengfei Wan 0001, Haotian Wang 0001, Wenjing Yang 0002 |
NeurIPS | 7 |
| 2025 | SeCon-RAG: A Two-Stage Semantic Filtering and Conflict-Free Framework for Trustworthy RAGabstractRetrieval-augmented generation (RAG) systems enhance large language models (LLMs) with external knowledge but are vulnerable to corpus poisoning and contamination attacks, which can compromise output integrity. Existing defenses often apply aggressive filtering, leading to unnecessary loss of valuable information and reduced reliability in generation.
To address this problem, we propose a two-stage semantic filtering and conflict-free framework for trustworthy RAG.
In the first stage, we perform a joint filter with semantic and cluster-based filtering which is guided by the Entity-intent-relation extractor (EIRE). EIRE extracts entities, latent objectives, and entity relations from both the user query and filtered documents, scores their semantic relevance, and selectively adds valuable documents into the clean retrieval database.
In the second stage, we proposed an EIRE-guided conflict-aware filtering module, which analyzes semantic consistency between the query, candidate answers, and retrieved knowledge before final answer generation, filtering out internal and external contradictions that could mislead the model.
Through this two-stage process, SeCon-RAG effectively preserves useful knowledge while mitigating conflict contamination, achieving significant improvements in both generation robustness and output trustworthiness.
Extensive experiments across various LLMs and datasets demonstrate that the proposed SeCon-RAG markedly outperforms state-of-the-art defense methods. Xiaonan Si, Meilin Zhu, Simeng Qin, Lijia Yu, Shuaitong Liu, Xinfeng Li, Ranjie Duan, Xiaojun Jia |
NeurIPS | 7 |
| 2025 | LIFEBENCH: Evaluating Length Instruction Following in Large Language ModelsabstractWhile large language models (LLMs) can solve PhD-level reasoning problems over long context inputs, they still struggle with a seemingly simpler task: following explicit length instructions—e.g., write a 10,000-word novel. Additionally, models often generate far too short outputs, terminate prematurely, or even refuse the request. Existing benchmarks focus primarily on evaluating generations quality, but often overlook whether the generations meet length constraints. To this end, we introduce Length Instruction Following Evaluation Benchmark (LIFEBench) to comprehensively evaluate LLMs' ability to follow length instructions across diverse tasks and a wide range of specified lengths. LIFEBench consists of 10,800 instances across 4 task categories in both English and Chinese, covering length constraints ranging from 16 to 8192 words. We evaluate 26 widely-used LLMs and find that most models reasonably follow short-length instructions but deteriorate sharply beyond a certain threshold. Surprisingly, almost all models fail to reach the vendor-claimed maximum output lengths in practice, as further confirmed by our evaluations extending up to 32K words. Even long-context LLMs, despite their extended input-output windows, counterintuitively fail to improve length-instructions following. Notably, Reasoning LLMs outperform even specialized long-text generation models, achieving state-of-the-art length following. Overall, LIFEBench uncovers fundamental limitations in current LLMs' length instructions following ability, offering critical insights for future progress. Zhenhong Zhou, Kun Wang 0056, Junfeng Fang, Rongwu Xu, Yuanhe Zhang, Xinfeng Li, Li Sun 0008, Lingjuan Lyu, Sen Su |
NeurIPS | 9 |
| 2025 | PATFinger: Prompt-Adapted Transferable Fingerprinting against Unauthorized Multimodal Dataset UsageabstractThe multimodal datasets can be leveraged to pre-train large-scale vision-language models by providing cross-modal semantics. Current endeavors for determining the usage of datasets mainly focus on single-modal dataset ownership verification through intrusive methods and non-intrusive techniques, while cross-modal approaches remain under-explored. Intrusive methods can adapt to multimodal datasets but degrade model accuracy, while non-intrusive methods rely on label-driven decision boundaries that fail to guarantee stable behaviors for verification. To address these issues, we propose a novel prompt-adapted transferable fingerprinting scheme from a training-free perspective, called PATFinger, which incorporates the global optimal perturbation (GOP) and the adaptive prompts to capture dataset-specific distribution characteristics. Our scheme utilizes inherent dataset attributes as fingerprints instead of compelling the model to learn triggers. The GOP is derived from the sample distribution to maximize embedding drifts between different modalities. Subsequently, our PATFinger re-aligns the adaptive prompt with GOP samples to capture the cross-modal interactions on the carefully crafted surrogate model. This allows the dataset owner to check the usage of datasets by observing specific prediction behaviors linked to the PATFinger during retrieval queries. Extensive experiments demonstrate the effectiveness of our scheme against unauthorized multimodal dataset usage on various cross-modal retrieval architectures by 30% over state-of-the-art baselines. Ju Jia, Xiaojun Jia, Yihao Huang 0001, Xinfeng Li, Cong Wu 0003, Lina Wang 0001 |
SIGIR | 5 |
| 2025 | Neural Invisibility Cloak: Concealing Adversary in Images via Compromised AI-driven Image Signal Processing
Xiaoyu Ji 0001, Xinfeng Li, Ruoyan Xu, Wenyuan Xu 0001 |
USENIX Security Symposium | 3 |
| 2024 | SafeEar: Content Privacy-Preserving Audio Deepfake Detection
Xinfeng Li, Kai Li 0047, Yifan Zheng 0001, Chen Yan 0001, Xiaoyu Ji 0001, Wenyuan Xu 0001 |
CCS | 1 |
| 2024 | SafeGen: Mitigating Sexually Explicit Content Generation in Text-to-Image ModelsabstractText-to-image (T2I) models, such as Stable Diffusion, have exhibited remarkable performance in generating high-quality images from text descriptions in recent years. However, text-to-image models may be tricked into generating not-safe-for-work (NSFW) content, particularly in sexually explicit scenarios. Existing countermeasures mostly focus on filtering inappropriate inputs and outputs, or suppressing improper text embeddings, which can block sexually explicit content (e.g., naked) but may still be vulnerable to adversarial prompts -- inputs that appear innocent but are ill-intended. In this paper, we present SafeGen, a framework to mitigate sexual content generation by text-to-image models in a text-agnostic manner. The key idea is to eliminate explicit visual representations from the model regardless of the text input. In this way, the text-to-image model is resistant to adversarial prompts since such unsafe visual representations are obstructed from within. Extensive experiments conducted on four datasets and large-scale user studies demonstrate SafeGen's effectiveness in mitigating sexually explicit content generation while preserving the high-fidelity of benign images. SafeGen outperforms eight state-of-the-art baseline methods and achieves 99.4% sexual content removal performance. Furthermore, our constructed benchmark of adversarial prompts provides a basis for future development and evaluation of anti-NSFW-generation methods. Xinfeng Li, Jiangyi Deng, Chen Yan 0001, Yanjiao Chen, Xiaoyu Ji 0001, Wenyuan Xu 0001 |
CCS | 1 |
| 2024 | Legilimens: Practical and Unified Content Moderation for Large Language Model ServicesabstractGiven the societal impact of unsafe content generated by large language models (LLMs), ensuring that LLM services comply with safety standards is a crucial concern for LLM service providers. Common content moderation methods are limited by an effectiveness-and-efficiency dilemma, where simple models are fragile while sophisticated models consume excessive computational resources. In this paper, we reveal for the first time that effective and efficient content moderation can be achieved by extracting conceptual features from chat-oriented LLMs, despite their initial fine-tuning for conversation rather than content moderation. We propose a practical and unified content moderation framework for LLM services, named Legilimens, which features both effectiveness and efficiency. Our red-team model-based data augmentation enhances the robustness of Legilimens against state-of-the-art jailbreaking. Additionally, we develop a framework to theoretically analyze the cost-effectiveness of Legilimens compared to other methods Jialin Wu 0001, Jiangyi Deng, Shengyuan Pang, Yanjiao Chen, Xinfeng Li, Wenyuan Xu 0001 |
CCS | 6 |
| 2024 | Beyond Universal Transformer: Block Reusing with Adaptor in Transformer for Automatic Speech Recognition
Zhaoyi Liu 0003, Chang Zeng, Xinfeng Li |
ISNN | 4 |
| 2024 | Inaudible Adversarial Perturbation: Manipulating the Recognition of User Speech in Real Time
Xinfeng Li, Chen Yan 0001, Xuancun Lu, Xiaoyu Ji 0001, Wenyuan Xu 0001 |
NDSS | 1 |
| 2024 | Enrollment-Stage Backdoor Attacks on Speaker Recognition Systems via Adversarial UltrasoundabstractAutomatic Speaker Recognition Systems (SRSs) have been widely used in voice applications for personal identification and access control. A typical SRS consists of three stages, i.e., training, enrollment, and recognition. Previous work has revealed that SRSs can be bypassed by backdoor attacks at the training stage or by adversarial example attacks at the recognition stage. In this paper, we propose TUNER, a new type of backdoor attack against the enrollment stage of SRS via adversarial ultrasound modulation, which is inaudible, synchronization-free, content-independent, and black-box. Our key idea is to first inject the backdoor into the SRS with modulated ultrasound when a legitimate user initiates the enrollment, and afterward, the polluted SRS will grant access to both the legitimate user and the adversary with high confidence. Our attack faces a major challenge of unpredictable user articulation at the enrollment stage. To overcome this challenge, we generate the ultrasonic backdoor by augmenting the optimization process with random speech content, vocalizing time, and volume of the user. Furthermore, to achieve real-world robustness, we improve the ultrasonic signal over traditional methods using sparse frequency points, pre-compensation, and single-sideband (SSB) modulation. We extensively evaluate TUNER on two common datasets and seven representative SRS models, as well as its robustness against seven kinds of defenses. Results show that our attack can successfully bypass speaker recognition systems while remaining effective to various speakers, speech content, etc. To mitigate this newly discovered threat, we also provide discussions on potential countermeasures, limitations, and future works of this new threat. Xinfeng Li, Junning Ze, Chen Yan 0001, Yushi Cheng, Xiaoyu Ji 0001, Wenyuan Xu 0001 |
IEEE Internet Things J. | 1 |
| 2024 | Toward Pitch-Insensitive Speaker Verification via SoundfieldabstractAutomatic speaker verification systems (ASVs) verify a person’s identity by his/her voice and have been widely deployed for user authentication. However, existing ASVs are based on traditional audio spectral features and hence, perform poorly in verifying pitch-changed utterances from speakers with cold or sore throat. In this article, we propose soundfield tracker(SOFTER), a soundfield-based speaker verification system that can verify speakers regardless of the pitch changes.SOFTERis based on the observation that soundfield features reflect the speaker’s vocal tract, mouth, head, torso, etc., which are less affected by the pitch changes in speech signals.SOFTERcan be integrated into off-the-shelf smartphones without any hardware modifications. One major challenge is that the soundfield is sensitive to the distance between the speaker and the phone. To solve this problem, we propose a two-stage mechanism combining distance sensing and soundfield reconstruction, which enables to reconstruct the soundfield to a setting similar to the one in the enrollment phase, thus, the speaker can be verified from any distance to the phone. We compareSOFTERwith six state-of-the-art academic and commercial ASVs on two data sets of 134 speakers and 31000 speech samples. Results show thatSOFTERhas an equal error rate (EER) of 2.18% and 1.61% on the two data sets, respectively. Moreover,SOFTERoutperforms other ASVs by at least 24.67% on average in verifying pitch-varying or pathological speech samples, denoting an evidence ofSOFTER’s effectiveness in both normal and unhealthy user conditions. Xinfeng Li, Zhicong Zheng, Chen Yan 0001, Chaohao Li, Xiaoyu Ji 0001, Wenyuan Xu 0001 |
IEEE Internet Things J. | 1 |
| 2024 | Scoring Metrics of Assessing Voiceprint Distinctiveness Based on Speech Content and RateabstractA voiceprint is the distinctive pattern of human voices widely used for authentication in voice assistants. This paper investigates the impact of speech contents and speech rates on the distinctiveness of voiceprint, and has obtained answers to three questions by studying 2457 speakers and 21,500,000 test samples: 1) What are the influential factors that users can control to affect the distinctiveness of voiceprints? 2) How to quantify the distinctiveness for given speeches, e.g., the speech of wake-up words when activating voice assistants? 3) How to help users select wake-up words and adjust the speech rate to improve distinctiveness levels? To answer those questions, we break down speeches into phones, and experimentally obtain the correlation between false recognition rates and the richness, order, length, and elements of the phones. Then, we define the PROLE Score that can reflect the voice distinctiveness, and evaluate 30 wake-up words of 19 commercial voice assistant products to provide recommendations on selecting secure voiceprint words. We also measure the correlation between false recognition rates and speech rates, and define the TER Score that reveals the distance of distinctiveness from the secure voiceprint, and it guides users to adjust their speech rate to a secure value. He Ruiwen, Yushi Cheng, Junning Ze, Xinfeng Li, Xiaoyu Ji 0001, Wenyuan Xu 0001 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2023 | The Silent Manipulator: A Practical and Inaudible Backdoor Attack against Speech Recognition SystemsabstractBackdoor Attacks have been shown to pose significant threats to automatic speech recognition systems (ASRs). Existing success largely assumes backdoor triggering in the digital domain, or the victim will not notice the presence of triggering sounds in the physical domain. However, in practical victim-present scenarios, the over-the-air distortion of the backdoor trigger and the victim awareness raised by its audibility may invalidate such attacks. In this paper, we propose SMA, an inaudible grey-box backdoor attack that can be generalized to real-world scenarios where victims are present by exploiting both the vulnerability of microphones and neural networks. Specifically, we utilize the nonlinear effects of microphones to inject an inaudible ultrasonic trigger. To accurately characterize the microphone response to the crafted ultrasound, we construct a novel nonlinear transfer function for effective optimization. We also design optimization objectives to ensure triggers' robustness in the physical world and transferability on unseen ASR models. In practice, SMA can bypass the microphone's built-in filters and human perception, activating the implanted trigger in the ASRs inaudibly, regardless of whether the user is speaking. Extensive experiments show that the attack success rate of SMA can reach nearly 100% in the digital domain and over 85% against most microphones in the physical domains by only poisoning about 0.5% of the training audio dataset. Moreover, our attack can resist typical defense countermeasures to backdoor attacks. Zhicong Zheng, Xinfeng Li, Chen Yan 0001, Xiaoyu Ji 0001, Wenyuan Xu 0001 |
ACM Multimedia | 2 |
| 2023 | Learning Normality is Enough: A Software-based Mitigation against Inaudible Voice Attacks
Xinfeng Li, Xiaoyu Ji 0001, Chen Yan 0001, Chaohao Li, Zhenning Zhang, Wenyuan Xu 0001 |
USENIX Security Symposium | 1 |
| 2022 | UltraBD: Backdoor Attack against Automatic Speaker Verification Systems via Adversarial UltrasoundabstractAutomatic speaker verification (ASV) systems have been widely applied in voice user interfaces to conduct person identification and access control via voiceprints. A typical ASV system consists of three stages, i.e., training, enrollment, and verification. Previous work has revealed that the ASV system can be bypassed at the training stage by backdoor attacks and at the verification stage by adversarial example attacks. In this paper, we propose a new type of backdoor attack aimed at the enrollment stage via adversarial ultrasound, named UltraBD, which is highly imperceptible, synchronization-free, and content-independent. By simultaneously injecting the ultrasound backdoor examples when the legitimate user initiates the enrollment, the polluted voiceprints stored in the ASV systems grant access to both the legitimate user and the adversary with relatively high confidence. Despite the challenges, i.e., when, what, and how the legitimate user articulates at the enrollment stage can be remarkably unpredictable and various, we managed to launch UltraBD by augmenting the generation and optimization process of the ultrasound backdoor examples with the randomness of synchronous time and relative amplitude ratio. Furthermore, we optimize the modulation mechanism of adversarial ultrasound by tuning the baseband signal on limited signal frequency points to improve its robustness in the physical world setting. We validate UltraBD on two common datasets together with two open-source ASV models. Results show that UltraBD can be robust to various configurations, e.g., different speakers and utterance content. In sum, our attack calls attention to a new attack surface of ASV systems and sheds light on its fundamental mechanisms. Junning Ze, Xinfeng Li, Yushi Cheng, Xiaoyu Ji 0001, Wenyuan Xu 0001 |
ICPADS | 2 |
| 2022 | "OK, Siri" or "Hey, Google": Evaluating Voiceprint Distinctiveness via Content-based PROLE Score
He Ruiwen, Xiaoyu Ji 0001, Xinfeng Li, Yushi Cheng, Wenyuan Xu 0001 |
USENIX Security Symposium | 3 |
| 2021 | Integrating Real-Time Entity Resolution with Top-N Join Query Processing
Xinfeng Li, Yonggang Wei, Qin Ma 0006, Weiyi Meng |
KSEM | 2 |
| 2021 | EarArray: Defending against DolphinAttack via Acoustic Attenuation
Xiaoyu Ji 0001, Xinfeng Li, Gang Qu 0001, Wenyuan Xu 0001 |
NDSS | 3 |
| 2020 | Prediction of Weather Radar Images via a Deep LSTM for NowcastingabstractWeather radar images provide critical information for mesoscale weather nowcasting which plays significant roles in a range of fields including civil aviation and navigation. Differed from traditional radar exploration methods, this paper presents a novel prediction model based on a deep recurrent neural network (DeepRNN). The approach converts the task of nowcasting to a task of image series prediction. We first design a new loss function that pays more attention to the changes of images in the input sequence. In the mean while, an image discriminator is incorporated into the model to improve the visual quality of predicted images. Furthermore, optical flow is explored to preserve the the motion information. The prediction results are evaluated based on widely used statistic scores. The experimental results show that the proposed model leads to significant improvement in tasks of 2 hours forecasting of radar echo. Guang Yao, Zongxuan Liu, Xufeng Guo, Chaoshi Wei, Xinfeng Li |
IJCNN | 5 |
| 2017 | EV-Matching: Bridging Large Visual Data and Electronic Data for Efficient SurveillanceabstractVisual (V) surveillance systems are extensively deployed and becoming the largest source of big data. On the other hand, electronic (E) data also plays an important role in surveillance and its amount increases explosively with the ubiquity of mobile devices. One of the major problems in surveillance is to determine human objects' identities among different surveillance scenes. Traditional way of processing big V and E datasets separately does not serve the purpose well because V data and E data are imperfect alone for information gathering and retrieval. Matching human objects in the two datasets can merge the good of the two for efficient large-scale surveillance. Yet such matching across two heterogeneous big datasets is challenging. In this paper, we propose an efficient set of parallel algorithms, called EV-Matching, to bridge big E and V data. We match E and V data based on their spatiotemporal correlation. The EV-Matching algorithms are implemented on Apache Spark to further accelerate the whole procedure. We conduct extensive experiments on a large synthetic dataset under different settings. Results demonstrate the feasibility and efficiency of our proposed algorithms. Fan Yang 0059, Guoxing Chen, Qiang Zhai, Xinfeng Li, Jin Teng, Junda Zhu 0001, Dong Xuan, Biao Chen 0002, Wei Zhao 0001 |
ICDCS | 5 |
| 2017 | SurvSurf: human retrieval on large surveillance video data
Sihao Ding 0001, Ying Li 0138, Xinfeng Li, Qiang Zhai, Adam C. Champion, Junda Zhu 0001, Dong Xuan, Yuan F. Zheng |
Multim. Tools Appl. | 4 |
| 2017 | Traffic At-a-Glance: Time-Bounded Analytics on Large Visual Traffic DataabstractMassive visual traffic data have become available recently. Though it opens the realm of intelligent traffic analysis, processing the data in a timely manner is difficult yet critical to time sensitive decisions, which are typical to traffic related management. In this paper, we study time-bounded aggregation analytics on large visual traffic data including traffic images and videos. We first find that current MapReduce framework can not work well due to two challenges: first, significant dual diversities exist on data distributions and processing time; second, apriori knowledge on these distributions and time costs are not always available. However, we also observe spatial and temporal locality on data values and processing time. Based on the examination, we design Traffic At-a-Glance (TaG), an augmented MapReduce framework for time-bounded traffic analytics jobs. Particularly, we propose a novel sampling algorithm that exploits traffic data localities and stratifies samples based on data distributions and processing time. It runs in an iterative, adaptive manner without apriori knowledge. Moreover, we propose a heuristic scheduling algorithm with considerations of batch processing overhead. Further, we refine the load balancing mechanism based on data processing time locality to respect job time bounds. In addition, we extend TaG to well handle traffic videos by sampling video data based on motion information encoded in the videos. We implement TaG on Hadoop and conduct extensive experiments on a large visual traffic dataset. The evaluations on different data sizes show TaG is able to achieve high accuracy within time bounds. Xinfeng Li, Fan Yang 0059, Jin Teng, Sihao Ding 0001, Yuan F. Zheng, Dong Xuan, Biao Chen 0002, Wei Zhao 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2016 | Traffic at-a-glance: Time-bounded analytics on large visual traffic dataabstractMassive visual traffic data have become available recently, which provides an opportunity for intelligent traffic analysis. Timely processing is particularly necessary for traffic analysis. In this paper, we study time-bounded aggregation analytics on large visual traffic data. We first find that current MapReduce framework can not work well due to two challenges: first, significant dual diversities exist on data distributions and processing time; second, no apriori knowledge on these distributions and time costs is available. However, we also observe spatial and temporal locality on data values and processing time. Based on the examination, we design TaG, an augmented MapReduce framework for time-bounded traffic analytics jobs. Particularly, we propose a novel sampling algorithm that exploits traffic data localities and stratifies samples based on data distributions and processing time. It runs in an iterative, adaptive manner without apriori knowledge. Moreover, we propose a heuristic scheduling algorithm with considerations of batch processing overhead. Further, we refine load balancing mechanism based on data processing time locality to respect job time bounds. We implement TaG on Hadoop and conduct extensive experiments on a large traffic image dataset. The evaluations on different data sizes show TaG is able to achieve high accuracy within different time bounds. Xinfeng Li, Fan Yang 0059, Jin Teng, Dong Xuan, Biao Chen 0002 |
INFOCOM | 1 |
| 2015 | VM-tracking: Visual-motion sensing integration for real-time human trackingabstractHuman tracking in video has many practical applications such as visual guided navigation, assisted living, etc. In such applications, it is necessary to accurately track multiple humans across multiple cameras, subject to real-time constraints. Despite recent advances in visual tracking research, the tracking systems purely relying on visual information fail to meet the accuracy and real-time requirements at the same time. In this paper, we present a novel accurate and real-time human tracking system called VM-Tracking. The system aggregates the information of motion (M) sensor on human, and integrates it with visual (V) data based on physical locations. The system has two key features, i.e. location-based VM fusion and appearance-free tracking, which significantly distinguish itself from other existing human tracking systems. We have implemented the VM-Tracking system and conducted comprehensive experiments on challenging scenarios. Qiang Zhai, Sihao Ding 0001, Xinfeng Li, Fan Yang 0059, Jin Teng, Junda Zhu 0001, Dong Xuan, Yuan F. Zheng, Wei Zhao 0001 |
INFOCOM | 3 |
| 2014 | E-Shadow: Lubricating Social Interaction Using Mobile PhonesabstractIn this paper, we propose E-Shadow, a distributed mobile phone-based local social networking system. E-Shadow has two main components: (1) Local profiles. They enable E-Shadow users to record and share their names, interests, and other information with fine-grained privacy controls. (2) Mobile phone based local social interaction tools. E-Shadow provides mobile phone software that enables rich social interactions. The software maps proximate users’ local profiles to their human owners and enables user communication and content sharing. We have designed and implemented E-Shadow on mobile phones. In our E-Shadow system, we allow users to perform dynamic and layered information publishing, making use of interpersonal relevance in space and time. Our system also provides a mechanism to help users perform direction-driven localization of an E-Shadow and match it with its owner. Experiments on real world Windows Mobile phones and large-scale simulations show that our system disseminates information efficiently and helps receivers find the direction of a specific E-Shadow with accuracy. We believe our E-Shadow concept and system can lead to a more tightly-knit temporary community in one’s physical vicinity. Jin Teng, Boying Zhang, Xinfeng Li, Xiaole Bai, Dong Xuan |
IEEE Trans. Computers | 3 |
| 2014 | TurfCast: A Service for Controlling Information Dissemination in Wireless NetworksabstractRecent years have witnessed mass proliferation of mobile devices with rich wireless communication capabilities as well as emerging mobile device-based information dissemination applications that leverage these capabilities. This paper proposes TurfCast, a novel information dissemination service that selectively broadcasts information in particular "turfs,â abstract logical spaces in which receivers are situated. Such turfs can be temporal or spatial based on receivers' lingering time or physical areas, respectively. TurfCast has many applications such as electronic proximity advertising and mobile social networking. To enable TurfCast, we propose two supporting technologies: TurfCode and TurfBurst. TurfCode is a nested 0-1 fountain code that enables the broadcaster to transmit either all information or none at all to receivers. TurfBurst exploits the Shannon bound to differentiate among receivers: those who cannot receive information fast enough receive none at all, even if they linger near the broadcaster. We implement TurfCast on real-world devices and conduct experiments in both indoor and outdoor environments. Our experimental results illustrate TurfCast's potential for controlling information dissemination in wireless networks. Xinfeng Li, Jin Teng, Boying Zhang, Adam C. Champion, Dong Xuan |
IEEE Trans. Mob. Comput. | 1 |
| 2014 | EV-Loc: Integrating Electronic and Visual Signals for Accurate LocalizationabstractNowadays, more and more objects can be represented with electronic identifiers, e.g., people can be recognized from their laptops' MACs, and products can be identified by their RFID numbers. Localizing electronic identifiers is more and more important for a fully digitalized life. However, traditional wireless localization techniques are not satisfactory in performance to determine these electronic identifiers' positions. Some of them require costly hardware to achieve high accuracy and, hence, are not practical. The others are inaccurate and not robust against environmental noises, e.g., RSSI-based localization. Therefore, an accurate and practical approach for localizing electronic identifiers is needed. In this paper, we propose a new localization technique called EV-Loc. In EV-Loc, we make use of visual signals to help improve the accuracy of wireless localization. Our technique fully takes advantage of the high accuracy of visual signals and pervasiveness of electronic signals. To effectively couple these two signals together, we have designed an E-V match engine to find the correspondence between an object's electronic identifier and its visual appearance. We have implemented our technique on mobile devices and evaluated it in the real world. The localization error is less than 1 m. We have also evaluated our approach using large-scale simulations. The results show that our approach is accurate and robust. Jin Teng, Boying Zhang, Junda Zhu 0001, Xinfeng Li, Dong Xuan, Yuan F. Zheng |
IEEE/ACM Trans. Netw. | 4 |
| 2013 | D2Taint: Differentiated and dynamic information flow tracking on smartphones for numerous data sourcesabstractWith smartphones' meteoric growth in recent years, leaking sensitive information from them has become an increasingly critical issue. Such sensitive information can originate from smartphones themselves (e.g., location information) or from many Internet sources (e.g., bank accounts, emails). While prior work has demonstrated information flow tracking's (IFT's) effectiveness at detecting information leakage from smartphones, it can only handle a limited number of sensitive information sources. This paper presents a novel IFT tagging strategy using differentiated and dynamic tagging. We partition information sources into differentiated classes and store them in fixed-length tags. We adjust tag structure based on time-varying received information sources. Our tagging strategy enables us to track at runtime numerous information sources in multiple classes and rapidly detect information leakage from any of these sources. We design and implement D2Taint, an IFT system using our tagging strategy on real-world smartphones. We experimentally evaluate D2Taint's effectiveness with 84 real-world applications downloaded from Google Play. D2Taint reports that over 80% of them leak data to third-party destinations; 14% leak highly sensitive data. Our experimental evaluation using a standard benchmark tool illustrates D2Taint's effectiveness at handling many information sources on smartphones with moderate runtime and space overhead. Boxuan Gu, Xinfeng Li, Adam C. Champion, Zhezhe Chen, Dong Xuan |
INFOCOM | 2 |
| 2013 | EV-Human: Human localization via visual estimation of body electronic interferenceabstractHuman localization is an enabling technology for many mobile applications. As more and more people carry mobile phones with them, we can now localize a person by localizing his mobile phone. However, it is observed that presence of human bodies introduces heavy interference to mobile phone signals. This has been one of the major causes of inaccurate wireless localization for humans. In this paper, we propose using video cameras to help estimate human body's interference on mobile device's signals. We combine human orientation detection and human/phone/AP relative position inference estimation to better measure how a human blocks or reflects wireless signals. We have also developed a signal distortion compensation model. Based on these technologies, we have implemented a human localization system called EV-Human. Real world experiments show that our EV-system can accurately and robustly localize humans. Xinfeng Li, Jin Teng, Qiang Zhai, Junda Zhu 0001, Dong Xuan, Yuan F. Zheng, Wei Zhao 0001 |
INFOCOM | 1 |
| 2013 | On wireless network coverage in bounded areasabstractIn this paper, we study the problem of wireless coverage in bounded areas. Coverage is one of the fundamental requirements of wireless networks. There has been considerable research on optimal coverage of infinitely large areas. However, in the real world, the deployment areas of wireless networks are always geographically bounded. It is a much more challenging and significant problem to find optimal deployment patterns to cover bounded areas. In this paper, we approach this problem starting from the development of tight lower bounds on the number of nodes needed to cover a bounded area. Then we design several deployment patterns for different kinds of convex and concave shapes such as rectangles and L-shapes. These patterns require only few more nodes than the theoretical lower bound, and can achieve efficient coverage. We have also carefully addressed and evaluated practical conditions such as coverage modeling and connectivity regarding our deployment patterns. Zuoming Yu, Jin Teng, Xinfeng Li, Dong Xuan |
INFOCOM | 3 |
| 2012 | TurfCast: A service for controlling information dissemination in wireless networksabstractRecent years have witnessed mass proliferation of mobile devices with rich wireless communication capabilities as well as emerging mobile device based information dissemination applications that leverage these capabilities. This paper proposes TurfCast, a novel information dissemination service that selectively broadcasts information in particular “turfs,” abstract logical spaces in which receivers are situated. Such turfs can be temporal or spatial based on receivers' lingering time or physical areas, respectively. TurfCast has many applications such as electronic proximity advertising and mobile social networking. To enable TurfCast, we propose two supporting technologies: TurfCode and TurfBurst. TurfCode is a nested 0-1 fountain code that enables the broadcaster to transmit either all information or none at all to receivers. TurfBurst exploits the Shannon bound to differentiate among receivers: those who cannot receive information fast enough receive none at all, even if they linger near the broadcaster. We implement TurfCast on real-world devices and conduct experiments in both indoor and outdoor environments. Our experimental results illustrate TurfCast's potential for controlling information dissemination in wireless networks. Xinfeng Li, Jin Teng, Boying Zhang, Adam C. Champion, Dong Xuan |
INFOCOM | 1 |
| 2012 | EV-Loc: integrating electronic and visual signals for accurate localizationabstractNowadays, an increasing number of objects can be represented by their wireless electronic identifiers. For example, people can be recognized by their phone numbers or their phones' WiFi' MAC addresses and products can be identified by their RFID numbers. Localizing objects with electronic identifiers is increasingly important as our lives become increasingly "digitalized". However, traditional wireless localization techniques cannot meet the fast growing needs of accurate and cost efficient localization. Some of these techniques require expensive hardware to achieve high accuracy, which is impractical for massive deployment. Others, such as WiFi RSSI based localization, are inaccurate and not robust to environmental noise. In this paper, we propose a new localization technique called EV-Loc. In EV-Loc, we use visual signals to help improve the accuracy of wireless localization. Our technique fully leverages visual signals' high accuracy and electronic signals' pervasiveness. To effectively couple these two signals, we design an E-V match engine to find the correspondence between an object's electronic identifier and its visual appearance. We implement our technique on mobile devices and evaluate it in real-world scenarios. The localization error is less than 1 m. We also evaluate our approach using large scale simulations. The results show that our approach is accurate and robust. Boying Zhang, Jin Teng, Junda Zhu 0001, Xinfeng Li, Dong Xuan, Yuan F. Zheng |
MobiHoc | 4 |
| 2012 | Enclave: Promoting Unobtrusive and Secure Mobile Communications with a Ubiquitous Electronic World
Adam C. Champion, Xinfeng Li, Qiang Zhai, Jin Teng, Dong Xuan |
WASA | 2 |
| 2011 | E-Shadow: Lubricating Social Interaction Using Mobile PhonesabstractIn this paper, we propose E-Shadow, a distributed mobile phone-based local social networking system. E-Shadow has two main components: (1) Local profiles. They enable EShadow users to record and share their names, interests, and other information with fine-grained privacy controls. (2) Mobile phone based local social interaction tools. E-Shadow provides mobile phone software that enables rich social interactions. The software maps proximate users' local profiles to their human owners and enables user communication and content sharing. We have designed and implemented E-Shadow on mobile phones. In our E-Shadow system, we allow users to perform dynamic and layered information publishing, making use of interpersonal relevance. Our system also provides a mechanism to help users perform direction-driven localization of an E-Shadow and match it with its owner. Experiments on real world Windows Mobile phones and large-scale simulations show that our system disseminates information efficiently and helps receivers find the direction of a specific E-Shadow with accuracy. We believe our E-Shadow concept and system can lead to a more tightly-knit temporary community in one's physical vicinity. Jin Teng, Boying Zhang, Xinfeng Li, Xiaole Bai, Dong Xuan |
ICDCS | 3 |
| 2009 | Enhanced Location Privacy Protection of Base Station in Wireless Sensor NetworksabstractLocation privacy in wireless sensor networks has gained a wide concern. Particularly, the location privacy of base station requires ultimate protection due to its crucial position in wireless sensor networks. In this paper, we propose an efficient scheme, consisting of anonymous topology discovery and intelligent fake packet injection (IFPI), to protect the location privacy of base station. Anonymous topology discovery eliminates the potential threats against base station within topology discovery period. On the other hand, IFPI enhances privacy protection strength during data transmission period. Under given conditions, comprehensive simulations demonstrate that our scheme significantly improves privacy strength compared with existing strategies. Xinfeng Li, Zhiguo Wan, Ming Gu 0001 |
MSN | 1 |
| 2009 | CLEAR: A confidential and Lifetime-Aware Routing Protocol for wireless sensor networkabstractA key challenge of the resource-constrained wireless sensor network is to prolong the lifetime as long as possible. Researchers have proved data transmission consumes most energy of sensor nodes, and the routing policy has great influence on network lifetime. Another important factor is the location confidentiality of base station that may be easily captured by a packet-tracing adversary due to the open communication nature of sensor network. In this paper, we design a new routing scheme called CLEAR: A Confidential and Lifetime-Aware Routing Protocol. In CLEAR, the paths between sources and base station keep changing throughout the whole lifetime in order to make best use of the limited node energy. On the other hand, we introduce an extended confidentiality protection mechanism called Branching, combining with the basic confidentiality feature of CLEAR as a whole. Comprehensive simulations prove that CLEAR extends the lifetime of sensor network almost twice as much as that of the basic routing protocol, and prominently enhances the base station location confidentiality compared with several existing schemes. Xinfeng Li, Zhiguo Wan, Ming Gu 0001 |
PIMRC | 2 |