Xianjun Yang

dblp:37/10237 · DBLP profile ↗
← Back
34ranked-venue papers
10as first author
24since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 3 first-author · 16 since 2021Computer networks · 7 · 6 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Hallucination as a Computational Boundary: A Hierarchy of Inevitability and the Oracle Escape
abstract
The illusion phenomenon of large language models (LLMs) is the core obstacle to their reliable deployment. This article formalizes the large language model as a probabilistic Turing machine by constructing a "computational necessity hierarchy", and for the first time proves the illusions are inevitable on diagonalization, incomputability, and information theory boundaries supported by the new "learner pump lemma". However, we propose two "escape routes": one is to model Retrieval Enhanced Generations (RAGs) as oracle machines, proving their absolute escape through "computational jumps", providing the first formal theory for the effectiveness of RAGs; The second is to formalize continuous learning as an "internalized oracle" mechanism and implement this path through a novel neural game theory framework.Finally, this article proposes a feasible new principle for artificial intelligence security - Computational Class Alignment (CCA), which requires strict matching between task complexity and the actual computing power of the system, providing theoretical support for the secure application of artificial intelligence.
Zenghui Ding, Jianqing Gao, Xianjun Yang
AAAI5
2025 Unveiling the Impact of Coding Data Instruction Fine-Tuning on Large Language Models Reasoning
abstract
Instruction Fine-Tuning (IFT) significantly enhances the zero-shot capabilities of pretrained Large Language Models (LLMs). While coding data is known to boost LLM reasoning abilities during pretraining, its role in activating internal reasoning capacities during IFT remains understudied. This paper investigates a key question: How does coding data impact LLMs' reasoning capacities during IFT stage? To explore this, we thoroughly examine the impact of coding data across different coding data proportions, model families, sizes, and reasoning domains, from various perspectives. Specifically, we create three IFT datasets with increasing coding data proportions, fine-tune six LLM backbones across different families and scales on these datasets, evaluate the tuned models' performance across twelve tasks in three reasoning domains, and analyze the outcomes from three broad-to-granular perspectives: overall, domain-level, and task-specific. Our holistic analysis provides valuable insights into each perspective. First, coding data tuning enhances the overall reasoning capabilities of LLMs across different model families and scales. Moreover, while the impact of coding data varies by domain, it shows consistent trends within each domain across different model families and scales. Additionally, coding data generally provides comparable task-specific benefits across model families, with optimal proportions in IFT datasets being task-dependent.
Xinlu Zhang, Zhiyu Chen 0002, Xi Ye 0003, Xianjun Yang, Lichang Chen, William Yang Wang, Linda R. Petzold
AAAI4
2025 Recognizing Diabetic Foot Progression via Multimodal Fusion of Infrared Thermography and Clinical Data
abstract
Diabetic foot (DF), a severe diabetes complication, remains a leading cause of lower-limb amputations, underscoring the urgent need for early-stage detection to enable timely intervention and improve outcomes. However, current approaches often fail to distinguish diabetic patients without foot complications (DM) from those with DF, as single-modality infrared thermography (IRT) struggles with subtle thermal cues and lacks localization of pathological regions. In this paper, we propose DFP-MMNet (Diabetic Foot Progression MultiModal Network), a novel two-stage multimodal framework that addresses this challenge. In the first stage, a deep registration model, guided by corresponding RGB images, accurately localizes lesion Regions of Interest (ROIs) on thermograms. In the second stage, a dual-branch network fuses deep visual features from the ROI (via RegNetY-16GF) with semantic embeddings derived from structured clinical data, encoded using CLIP text branch. Extensive experiments on the collected ITC dataset demonstrate that DFP-MMNet achieves an F1-score of 82.7%, outperforming unimodal baselines by over 15%, offering a robust and accurate solution for early DF recognition.
Huarui Liu, Qiankun Li 0004, Xianjun Yang, Yining Sun
BIBM5
2025 Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding
abstract
Recent advancements in multimodal large language models (MLLMs) have opened new avenues for video understanding. However, achieving high fidelity in zero-shot video tasks remains challenging. Traditional video processing methods rely heavily on fine-tuning to capture nuanced spatial-temporal details, which incurs significant data and computation costs. In contrast, training-free approaches, though efficient, often lack robustness in preserving context-rich features across complex video content. To this end, we propose DYTO, a novel dynamic token merging framework for zero-shot video understanding that adaptively optimizes token efficiency while preserving crucial scene details. DYTO integrates a hierarchical frame selection and a bipartite token merging strategy to dynamically cluster key frames and selectively compress token sequences, striking a balance between computational efficiency with semantic richness. Extensive experiments across multiple benchmarks demonstrate the effectiveness of DYTO, achieving superior performance compared to both fine-tuned and training-free methods and setting a new state-of-the-art for zero-shot video understanding.
Zhuokai Zhao, Zhaorun Chen, Zenghui Ding, Xianjun Yang, Yining Sun
ICCV5
2025 Weak-to-Strong Jailbreaking on Large Language Models
abstract
Large language models (LLMs) are vulnerable to jailbreak attacks -- resulting in harmful, unethical, or biased text generations. However, existing jailbreaking methods are computationally costly. In this paper, we propose the **weak-to-strong** jailbreaking attack, an efficient inference time attack for aligned LLMs to produce harmful text. Our key intuition is based on the observation that jailbroken and aligned models only differ in their initial decoding distributions. The weak-to-strong attack's key technical insight is using two smaller models (a safe and an unsafe one) to adversarially modify a significantly larger safe model's decoding probabilities. We evaluate the weak-to-strong attack on 5 diverse open-source LLMs from 3 organizations. The results show our method can increase the misalignment rate to over 99\% on two datasets with just one forward pass per example. Our study exposes an urgent safety issue that needs to be addressed when aligning LLMs. As an initial attempt, we propose a defense strategy to protect against such attacks, but creating more advanced defenses remains challenging. The code for replicating the method is available at https://github.com/XuandongZhao/weak-to-strong.
Xuandong Zhao, Xianjun Yang, Tianyu Pang, Lei Li 0005, Yu-Xiang Wang 0003, William Yang Wang
ICML2
2025 MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI Agents
abstract
Recent research has explored that LLM agents are vulnerable to indirect prompt injection (IPI) attacks, where malicious tasks embedded in tool-retrieved information can redirect the agent to take unauthorized actions. Existing defenses against IPI have significant limitations: either require essential model training resources, lack effectiveness against sophisticated attacks, or harm the normal utilities. We present MELON (Masked re-Execution and TooL comparisON), a novel IPI defense. Our approach builds on the observation that under a successful attack, the agent’s next action becomes less dependent on user tasks and more on malicious tasks. Following this, we design MELON to detect attacks by re-executing the agent’s trajectory with a masked user prompt modified through a masking function. We identify an attack if the actions generated in the original and masked executions are similar. We also include three key designs to reduce the potential false positives and false negatives. Extensive evaluation on the IPI benchmark AgentDojo demonstrates that MELON outperforms SOTA defenses in both attack prevention and utility preservation. Moreover, we show that combining MELON with a SOTA prompt augmentation defense (denoted as MELON-Aug) further improves its performance. We also conduct a detailed ablation study to validate our key designs. Code is available at https://github.com/kaijiezhu11/MELON.
Kaijie Zhu, Xianjun Yang, Jindong Wang 0001, Wenbo Guo 0002, William Yang Wang
ICML2
2025 CBT-Bench: Evaluating Large Language Models on Assisting Cognitive Behavior Therapy
abstract
Mian Zhang, Xianjun Yang, Xinlu Zhang, Travis Labrum, Jamie C. Chiu, Shaun M. Eack, Fei Fang, William Yang Wang, Zhiyu Chen. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Xianjun Yang, Xinlu Zhang, Travis Labrum, Jamie C. Chiu, Shaun M. Eack, Fei Fang 0001, William Yang Wang, Zhiyu Chen 0002
NAACL (Long Papers)2
2025 Deep Learning based Unified CSI Feedback for TDD and FDD Massive MIMO Systems
abstract
Accurate channel state information (CSI) is essential for maximizing massive MIMO throughput. While time-division duplexing (TDD) systems exploit channel reciprocity for easier downlink CSI (DL-CSI) acquisition from uplink CSI (UL-CSI), moderate channel differences still arise due to environmental factors and mobility. Frequency-division duplexing (FDD) systems also exhibit some channel reciprocity, allowing both TDD and FDD to reduce CSI feedback overhead. For devices supporting both modes, a unified CSI feedback framework can lower cost and energy consumption. This paper proposes a deep learning method based on a cascaded convolutional long short-term memory network, consisting of a main prediction module and a TDD preprocessing module. By leveraging the correlation between TDD and FDD channels under the same spatiotemporal conditions, the proposed method enhances TDD prediction performance and achieves unified CSI feedback for both TDD and FDD systems. Experimental results show that the scheme improves the squared generalized cosine similarity by 5% in TDD mode and by 10% in FDD mode.
Mingyu Jia, Shaoli Kang, Yuqi Xue, Yekang Wang, Xianjun Yang
VTC2025-Fall5
2025 AI-Enhanced CSI Feedback via Exploiting Multi-user Shared Information in mMIMO Systems
abstract
In massive multiple-input multiple-output (mMIMO) frequency division duplex (FDD) systems, user equipment (UEs) must feed back the channel state information (CSI) to the base station (BS) over the uplink to enhance spectral efficiency. However, in the millimeter wave (mmWave) or higher frequency bands, the significant increase in the number of transmit antennas results in prohibitive feedback overhead. To tackle this issue, we propose a deep learning based multi-user CSI compression framework called CoTransNet. For users in a neighboring area, CSI can be decomposed into two components: one is shared by all users in the area and another is unique to each individual UE. By leveraging shared information, CoTransNet extracts correlations among channels and eliminates redundant feedback information. Experimental results indicate that employing the CoTransNet architecture leads to an average improvement of approximately 4.98% in squared generalized cosine similarity (SGCS) across diverse feedback budgets and a reduction of about 3.54 dB in normalized mean-squared error (NMSE).
Yekang Wang, Shanzhi Chen, Shaoli Kang, Xianjun Yang, Mingyu Jia, Yuqi Xue
VTC2025-Fall4
2025 Deep Learning-Based Joint Prediction and Compression with Optimal Prediction Interval for High-speed CSI Feedback
abstract
In 6G-oriented ultra-large-scale multiple-input multiple-output (MIMO) systems, accurate channel state information (CSI) is crucial in achieving extremely high data rates. To address the challenge of poor timeliness in CSI feedback for high-speed mobile scenarios, this paper proposes a joint CSI prediction and compression scheme for CSI feedback. Traditional methods improve timeliness by increasing the feedback frequency but suffer from additional overhead. The core innovation of this proposal is the integration of CSI prediction and compression functions on the user equipment (UE) side using deep learning methods. Considering channel estimation errors, the paper investigates the optimal prediction time interval to maximize the system’s channel capacity. ConvLSTM network is used to predict the future CSI when PDSCH is transmitted with historical CSIs, followed by a Transformer model that performs adaptive compression under dynamic channel conditions. Experimental results demonstrate that the proposed method achieves a 7.06% performance improvement over the eType II codebook-based feedback scheme. It ensures both timeliness and accuracy of CSI feedback in high-speed mobility scenarios while maintaining low feedback overhead.
Yuqi Xue, Shaoli Kang, Xianjun Yang, Mingyu Jia, Yekang Wang
VTC2025-Fall3
2024 Navigating the OverKill in Large Language Models
abstract
Chenyu Shi, Xiao Wang, Qiming Ge, Songyang Gao, Xianjun Yang, Tao Gui, Qi Zhang, Xuanjing Huang, Xun Zhao, Dahua Lin. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Chenyu Shi, Xiao Wang 0042, Qiming Ge, Songyang Gao, Xianjun Yang, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001, Dahua Lin
ACL (1)5
2024 Large Language Models Can Be Contextual Privacy Protection Learners
abstract
Yijia Xiao, Yiqiao Jin, Yushi Bai, Yue Wu, Xianjun Yang, Xiao Luo, Wenchao Yu, Xujiang Zhao, Yanchi Liu, Quanquan Gu, Haifeng Chen, Wei Wang, Wei Cheng. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Yijia Xiao, Yiqiao Jin, Yushi Bai, Xianjun Yang, Xiao Luo 0001, Wenchao Yu, Xujiang Zhao, Yanchi Liu, Quanquan Gu, Wei Wang 0010, Wei Cheng 0002
EMNLP5
2024 A Self-Supervised Pressure Map Human Keypoint Detection Approch: Optimizing Generalization and Computational Efficiency Across Datasets
abstract
In environments where RGB images are inadequate, pressure maps is a viable alternative, garnering scholarly attention. This study introduces a novel self-supervised pressure map keypoint detection (SPMKD) method, addressing the current gap in specialized designs for human keypoint extraction from pressure maps. Central to our contribution is the Encoder-Fuser-Decoder (EFD) model, which is a robust framework that integrates a lightweight encoder for precise human keypoint detection, a fuser for efficient gradient propagation, and a decoder that transforms human keypoints into reconstructed pressure maps. This structure is further enhanced by the Classification-to-Regression Weight Transfer (CRWT) method, which fine-tunes accuracy through initial classification task training. This innovation not only enhances human keypoint generalization without manual annotations but also showcases remarkable efficiency and generalization, evidenced by a reduction to only 5.96% in FLOPs and 1.11% in parameter count compared to the baseline methods. Code is accessible at SPMKD-52CB.
Chengzhang Yu, Xianjun Yang, Wenxia Bao, Shaonan Wang, Zhiming Yao
ICASSP2
2024 DNA-GPT: Divergent N-Gram Analysis for Training-Free Detection of GPT-Generated Text
abstract
Large language models (LLMs) have notably enhanced the fluency and diversity of machine-generated text. However, this progress also presents a significant challenge in detecting the origin of a given text, and current research on detection methods lags behind the rapid evolution of LLMs. Conventional training-based methods have limitations in flexibility, particularly when adapting to new domains, and they often lack explanatory power. To address this gap, we propose a novel training-free detection strategy called Divergent N-Gram Analysis (DNA-GPT). Given a text, we first truncate it in the middle and then use only the preceding portion as input to the LLMs to regenerate the new remaining parts. By analyzing the differences between the original and new remaining parts through N-gram analysis in black-box or probability divergence in white-box, we can clearly illustrate significant discrepancies between machine-generated and human-written text. We conducted extensive experiments on the most advanced LLMs from OpenAI, including text-davinci-003, GPT-3.5-turbo, and GPT-4, as well as open-source models such as GPT-NeoX-20B and LLaMa-13B. Results show that our zero-shot approach exhibits state-of-the-art performance in distinguishing between human and GPT-generated text on four English and one German dataset, outperforming OpenAI's own classifier, which is trained on millions of text. Additionally, our methods provide reasonable explanations and evidence to support our claim, which is a unique feature of explainable detection. Our method is also robust under the revised text attack and can additionally solve model sourcing.
Xianjun Yang, Wei Cheng 0002, Linda R. Petzold, William Yang Wang
ICLR1
2024 Enhancing Small Medical Learners with Privacy-preserving Contextual Prompting
abstract
Large language models (LLMs) demonstrate remarkable medical expertise, but data privacy concerns impede their direct use in healthcare environments. Although offering improved data privacy protection, domain-specific small language models (SLMs) often underperform LLMs, emphasizing the need for methods that reduce this performance gap while alleviating privacy concerns. In this paper, we present a simple yet effective method that harnesses LLMs' medical proficiency to boost SLM performance in medical tasks under $privacy-restricted$ scenarios. Specifically, we mitigate patient privacy issues by extracting keywords from medical data and prompting the LLM to generate a medical knowledge-intensive context by simulating clinicians' thought processes. This context serves as additional input for SLMs, augmenting their decision-making capabilities. Our method significantly enhances performance in both few-shot and full training settings across three medical knowledge-intensive tasks, achieving up to a 22.57% increase in absolute accuracy compared to SLM fine-tuning without context, and sets new state-of-the-art results in two medical tasks within privacy-restricted scenarios. Further out-of-domain testing and experiments in two general domain datasets showcase its generalizability and broad applicability.
Xinlu Zhang, Xianjun Yang, Chenxin Tian, Yao Qin 0001, Linda R. Petzold
ICLR3
2024 Position: A Safe Harbor for AI Evaluation and Red Teaming
abstract
Independent evaluation and red teaming are critical for identifying the risks posed by generative AI systems. However, the terms of service and enforcement strategies used by prominent AI companies to deter model misuse have disincentives on good faith safety evaluations. This causes some researchers to fear that conducting such research or releasing their findings will result in account suspensions or legal reprisal. Although some companies offer researcher access programs, they are an inadequate substitute for independent research access, as they have limited community representation, receive inadequate funding, and lack independence from corporate incentives. We propose that major generative AI developers commit to providing a legal and technical safe harbor, protecting public interest safety research and removing the threat of account suspensions or legal reprisal. These proposals emerged from our collective experience conducting safety, privacy, and trustworthiness research on generative AI systems, where norms and incentives could be better aligned with public interests, without exacerbating model misuse. We believe these commitments are a necessary step towards more inclusive and unimpeded community efforts to tackle the risks of generative AI.
Shayne Longpre, Sayash Kapoor, Kevin Klyman, Ashwin Ramaswami, Rishi Bommasani, Borhane Blili-Hamelin, Yangsibo Huang, Aviya Skowron, Suhas Kotha, Yi Zeng 0005, Weiyan Shi 0001, Xianjun Yang, Reid Southen, Alexander Robey, Patrick Chao, Diyi Yang, Ruoxi Jia 0001, Daniel Kang 0001, Alex Pentland, Arvind Narayanan, Percy Liang, Peter Henderson 0002
ICML13
2024 DALD: Improving Logits-based Detector without Logits from Black-box LLMs
abstract
The advent of Large Language Models (LLMs) has revolutionized text generation, producing outputs that closely mimic human writing. This blurring of lines between machine- and human-written text presents new challenges in distinguishing one from the other – a task further complicated by the frequent updates and closed nature of leading proprietary LLMs. Traditional logits-based detection methods leverage surrogate models for identifying LLM-generated content when the exact logits are unavailable from black-box LLMs. However, these methods grapple with the misalignment between the distributions of the surrogate and the often undisclosed target models, leading to performance degradation, particularly with the introduction of new, closed-source models. Furthermore, while current methodologies are generally effective when the source model is identified, they falter in scenarios where the model version remains unknown, or the test set comprises outputs from various source models. To address these limitations, we present \textbf{D}istribution-\textbf{A}ligned \textbf{L}LMs \textbf{D}etection (DALD), an innovative framework that redefines the state-of-the-art performance in black-box text detection even without logits from source LLMs. DALD is designed to align the surrogate model's distribution with that of unknown target LLMs, ensuring enhanced detection capability and resilience against rapid model iterations with minimal training investment. By leveraging corpus samples from publicly accessible outputs of advanced models such as ChatGPT, GPT-4 and Claude-3, DALD fine-tunes surrogate models to synchronize with unknown source model distributions effectively. Our approach achieves SOTA performance in black-box settings on different advanced closed-source and open-source models. The versatility of our method enriches widely adopted zero-shot detection frameworks (DetectGPT, DNA-GPT, Fast-DetectGPT) with a `plug-and-play' enhancement feature. Extensive experiments validate that our methodology reliably secures high detection precision for LLM-generated text and effectively detects text from diverse model origins through a singular detector. Our method is also robust under the revised text attack and non-English texts.
Cong Zeng, Shengkun Tang, Xianjun Yang, Yuanzhou Chen, Yiyou Sun, Yao Li 0015, Wei Cheng 0002, Dongkuan Xu
NeurIPS3
2024 IAE-SDNet: An End-to-End Image Adaptive Enhancement and Wheat Scab Detection Network Using UAV
abstract
Wheat scab is a severe fungal disease caused by Fusarium. The precise detection of scab is crucial for improving the wheat yield. The utilization of RGB images captured by unmanned aerial vehicle (UAV) can effectively enhance the wheat scab detection efficiency. However, UAV images are influenced by environmental conditions. In addition, the scab in the image is extremely small and the background is complex. Therefore, to boost the wheat scab detection accuracy, this study proposes an end-to-end image adaptive enhancement and wheat scab detection network (IAE-SDNet). In the head of the network, an image adaptive enhancement (IAE) module aims to enhance the quality of UAV images and performs end-to-end learning with the transformer-based object detection module SDNet. A cascading inverse residual module (CIRMB) is designed to boost the feature extraction capacity for the disease area. A deformable multihead attention encoder (DMHA-Encoder) is deployed to augment the semantic information of the advanced features of wheat while maintaining the representation ability of global attention. A multiscale feature fusion block (MFFB) is designed to effectively fuse wheat disease features of different scales, reducing the impact caused by the inconsistency in size between the complex field background and the disease area. Finally, to optimize the detection of small lesion areas, an MNIoU bounding-box regression loss function is proposed. The experiment indicates that this method has improved average precision (AP) by 6% and the Recall by 5% compared with the baseline network. This study provides a feasible approach for the precise detection of wheat scab, with high application prospects.
Wenxia Bao, Penfei Zhang, Gensheng Hu, Linsheng Huang, Xianjun Yang
IEEE Trans. Geosci. Remote. Sens.6
2023 Few-Shot Document-Level Event Argument Extraction
abstract
Event argument extraction (EAE) has been well studied at the sentence level but under-explored at the document level.In this paper, we study to capture event arguments that actually spread across sentences in documents.Prior works usually assume full access to rich document supervision, ignoring the fact that the available argument annotation is usually limited.To fill this gap, we present FewDocAE, a Few-Shot Document-Level Event Argument Extraction benchmark, based on the existing documentlevel event extraction dataset.We first define the new problem and reconstruct the corpus by a novel N -Way-D-Doc sampling instead of the traditional N -Way-K-Shot strategy.Then we adjust the current document-level neural models into the few-shot setting to provide baseline results under in-and cross-domain settings.Since the argument extraction depends on the context from multiple sentences and the learning process is limited to very few examples, we find this novel task to be very challenging with substantively low performance.Considering FewDocAE is closely related to practical use under low-resource regimes, we hope this benchmark encourages more research in this direction.Our data and codes will be available online 1 .
Xianjun Yang, Linda R. Petzold
ACL (1)1
2023 ReDi: Efficient Learning-Free Diffusion Inference via Trajectory Retrieval
abstract
Diffusion models show promising generation capability for a variety of data. Despite their high generation quality, the inference for diffusion models is still time-consuming due to the numerous sampling iterations required. To accelerate the inference, we propose ReDi, a simple yet learning-free Retrieval-based Diffusion sampling framework. From a precomputed knowledge base, ReDi retrieves a trajectory similar to the partially generated trajectory at an early stage of generation, skips a large portion of intermediate steps, and continues sampling from a later step in the retrieved trajectory. We theoretically prove that the generation performance of ReDi is guaranteed. Our experiments demonstrate that ReDi improves the model inference efficiency by 2$\times$ speedup. Furthermore, ReDi is able to generalize well in zero-shot cross-domain image generation such as image stylization. The code and demo for ReDi is available at https://github.com/zkx06111/ReDiffusion.
Kexun Zhang, Xianjun Yang, William Yang Wang, Lei Li 0005
ICML2
2023 LLMScore: Unveiling the Power of Large Language Models in Text-to-Image Synthesis Evaluation
abstract
Existing automatic evaluation on text-to-image synthesis can only provide an image-text matching score, without considering the object-level compositionality, which results in poor correlation with human judgments. In this work, we propose LLMScore, a new framework that offers evaluation scores with multi-granularity compositionality. LLMScore leverages the large language models (LLMs) to evaluate text-to-image models. Initially, it transforms the image into image-level and object-level visual descriptions. Then an evaluation instruction is fed into the LLMs to measure the alignment between the synthesized image and the text, ultimately generating a score accompanied by a rationale. Our substantial analysis reveals the highest correlation of LLMScore with human judgments on a wide range of datasets (Attribute Binding Contrast, Concept Conjunction, MSCOCO, DrawBench, PaintSkills). Notably, our LLMScore achieves Kendall's tau correlation with human evaluations that is 58.8% and 31.2% higher than the commonly-used text-image matching metrics CLIP and BLIP, respectively.
Xianjun Yang, Xiujun Li, Xin Wang 0061, William Yang Wang
NeurIPS2
2022 Automatic Angle's classification based on the occlusal contact information
abstract
Malocclusion has a high prevalence in the population, which seriously affects patients’ oral and mental health. Angle’s classification is a widely accepted diagnostic standard for malocclusion, either requiring professional intervention and complicated procedures, or increasing radiation risks. This paper proposes a new method of Angle’s classification based on occlusal contact information to realize the automatic Angle’s classification. Firstly, a novel bite force measurement device is used to record the occlusal data of subjects with different occlusal categories, Meta-analysis evaluated several occlusion quantitative evaluation indicators. Then, the imbalance of the data set is improved by oversampling and popular machine learning models are used for training and performance evaluation. The result shows that the accuracy of the random forest model combined with occlusal contact information reaches 87.83%, and the performance of other evaluation indexes is good. It is demonstrated that machine learning models can be applied to Angle’s classification and shows the great potential of occlusal contact information in the aided diagnosis of oral diseases.
Zhiming Yao, Xianjun Yang, Yuanyin Wang, Wenhua Xu, Yining Sun
SMC4
2021 An Analysis of Relation Extraction within Sentences from Wet Lab Protocols
abstract
Wet lab protocols (WLPs) are sets of instructions written in domain-specific natural language for step-by-step biological experimental processes. There have been efforts to annotate WLPs for shallow semantic parsing to enable reproducible procedures, text mining, and automatic conversion into a machine-readable format. However, current methods have not fully exploited the relation extraction sub-task on the protocol corpus. Neural approaches have the potential to deal with the various noise and in-domain jargon in the texts. To explore the viability of neural methods for this task, we perform a thorough analysis of both graph and nongraph neural approaches. We find that both graph neural networks with generated parameters (GP-GNNs) and Context-Aware models show advantages in relation extraction and are well suited to our goal. Specifically, the GP-GNNs and Context-Aware models demonstrate similar performance on all three WLPs datasets when the full training set is used, both outperforming the previous best results significantly. This can be explained by the observation that considering multiple relations in a sentence enhances the predictive ability. In addition, our extensive experiments demonstrate that the Context-Aware approach in particular can achieve good results even with a limited amount of training data, providing new insights for low-resource scenarios.
Xianjun Yang, Xinlu Zhang, Julia Zuo, Stephen D. Wilson, Linda R. Petzold
IEEE BigData1
2021 Capturing Implicit Spatial Cues for Monocular 3d Hand Reconstruction
abstract
With the development of the parameterized hand model (e.g. MANO), it is possible to reconstruct the 3D hand mesh from a single 2D hand image by learning a few hand model parameters, rather than estimating hundreds of vertices on the mesh. However, it is highly non-linear to learn these parameters from the 2D hand image, as there is no explicit spatial correspondence between these parameters and image pixels. In this paper, we successfully leverage the graph convolutional network (GCN) to capture implicit spatial cues for fitting the well-known MANO hand model, thus greatly improving the performance of monocular 3D hand reconstruction. Our proposed MANO-GCN establishes the spatial hand mesh and hand joints graph for learning MANO parameters, with node features propagated along edges to utilize the spatial information. Among all monocular 3D hand reconstruction methods with MANO hand model, MANO-GCN achieves state-of-the-art accuracy on public FreiHAND and HO-3D benchmarks, without any bells and whistles. Code is available at https://github.com/ChenJoya/manogcn.
Joya Chen, Zhiming Yao, Xianjun Yang
ICME5
2017 Low Density Spreading Signature Vector Extension (LDS-SVE) for Uplink Multiple Access
abstract
Non-orthogonal multiple access (MA) is important for fulfilling 5G high connection density requirements in terms of accommodating more concurrent transmissions. Low density signature CDMA (LDS-CDMA) is a well-known non-orthogonal MA scheme where more users can be multiplexed thanks to low density spreading (LDS) and message passing algorithm (MPA). In this paper, a new LDS based MA scheme called LDS signature vector extension (LDS- SVE) is proposed for the uplink of orthogonal frequency division multiple (OFDM) systems. Compared with LDS-CDMA where each modulated symbol is individually spread onto multiple subcarriers to form a signature vector, LDS-SVE jointly transforms and spreads two modulated symbols onto twice the subcarriers to further exploit diversity which can be seen as a type of signature vector extension. The transformation is characterized by multiplying the modulated symbols with a transformation matrix which is optimized by minimizing the single-user raw bit error rate (BER). MPA with Turbo-equalization and successive interference cancellation (SIC) can be used at the receiver side to approach the single-user performance. Link-level simulations show potential performance benefit over OFDMA and LDS-CDMA.
Xianjun Yang
VTC Fall3
2016 Compressed sensing based ACK feedback for grant-free uplink data transmission in 5G mMTC
abstract
One of the three 5G key usage scenarios is massive machine type communication (mMTC), which is characterized by a very large number of connected devices typically transmitting a relatively low volume of non-delay sensitive data. To support the above mMTC communication, grant-free data transmission (GFDT) has been proposed, where mMTC devices have no dedicated transmission resource. One of the challenges faced by GFDT is to ensure the communication reliability, which can be enhanced by efficient acknowledgement (ACK) feedback. By exploiting the sparsity of user activity in mMTC, this paper proposes a compressed sensing based ACK feedback (CSAF) scheme for GFDT, to reduce the feedback signaling overhead. Furthermore, we propose a new CS recovery algorithm, named as regularized approximate message passing (RAMP), to improve the detection performance of CSAF in low SNR scenario. Simulation results show that, compared with the traditional method, the proposed CSAF scheme with the proposed RAMP algorithm reduces the feedback signaling overhead by up to 75.6% and has up to 3dB detection performance gain.
Xianjun Yang, Xin Wang 0133
PIMRC1
2015 Full-duplex and compressed sensing based neighbor discovery for wireless ad-hoc network
abstract
Neighbor discovery (ND) is an essential step to configure a wireless ad-hoc network, where a quick and accurate ND scheme is very important for subsequent network operations. However, the speeds of current ND schemes are still not quick enough and mainly restricted by the constraints of half-duplex and single packet reception (SPR). To overcome both of these two constraints, we propose a Full-duplex and Compressed sensing based Neighbor Discovery (FCND) scheme, which can complete the ND process in only one slot. Specifically, with the help of the emerging full-duplex technology, each node broadcasts a unique pseudo-random sequence, and receives a superposition of the sequences from its neighbors, simultaneously. Then, based on compressed sensing (CS) theory, each node decodes all its neighbor IDs from the received signal simultaneously, i.e., resolves the challenge of SPR. Mathematical analysis shows that decoding neighbor IDs in FCND scheme is actually a binary CS recovery problem. To further improve the decoding performance, we propose an adaptive thresholding and adjusting (ATA) algorithm. Simulation results show that, compared with existing schemes, the proposed FCND scheme can reduce the signaling overhead by up to 96%, and the proposed ATA algorithm can increase the decoding accuracy by up to 25%.
Xianjun Yang, Xin Wang 0133
WCNC1
2014 Anti-noise-folding regularized subspace pursuit recovery algorithm for noisy sparse signals
abstract
Denoising recovery algorithms are very important for the development of compressed sensing (CS) theory and its applications. Considering the noise present in both the original sparse signal x and the compressive measurements y, we propose a novel denoising recovery algorithm, named Regularized Subspace Pursuit (RSP). Firstly, by introducing a data pre-processing operation, the proposed algorithm alleviates the noise-folding effect caused by the noise added to x. Then, the indices of the nonzero elements in x are identified by regularizing the chosen columns of the measurement matrix. Afterwards, the chosen indices are updated by retaining only the largest entries in the Minimum Mean Square Error (MMSE) estimated signal. Simulation results show that, compared with the traditional orthogonal matching pursuit (OMP) algorithm, the proposed RSP algorithm increases the successful recovery rate (and reduces the reconstruction error) by up to 50% and 86% (35% and 65%) in high noise level scenarios and inadequate measurements scenarios, respectively.
Xianjun Yang, Qimei Cui, Eryk Dutkiewicz, Xiaojing Huang 0001, Xiaofeng Tao 0001, Gengfa Fang
WCNC1
2013 Analog compressed sensing for multiband signals with non-modulated Slepian basis
abstract
Recently, the recovery performance of analog Compressed Sensing (CS) has been significantly improved by representing multiband signals with the modulated and merged Slepian basis (MM-Slepian dictionary), which avoids the frequency leakage effect of the Discrete Fourier Transform (DFT) basis. However, the MM-Slepian dictionary has a very large scale and corresponds to a large-scale measurement matrix, which leads to high recovery computational complexity. This paper resolves the above problem by modulating and band-limiting the multiband signal rather than modulating the Slepian basis. Specifically, instead of using the MM-Slepian dictionary to represent the whole multiband signal, we propose to use the non-modulated Slepian basis to represent the modulated and band-limited version of the multiband signal based on the recently proposed Modulated Wideband Converter (MWC). Furthermore, based on the analytical derivation with the non-modulated Slepian basis, we propose an Interpolation Recovery (IR) algorithm to take full advantage of the Slepian basis, whereas the Direct Recovery (DR) algorithm using the Moore-Penrose pseudo-inverse cannot achieve this. Simulation results verify that, with low recovery computational load, the non-modulated Slepian basis combined with the IR algorithm improves the recovery SNR by up to 35 dB compared with the DFT basis in noise-free environment.
Xianjun Yang, Eryk Dutkiewicz, Qimei Cui, Xiaojing Huang 0001, Xiaofeng Tao 0001, Gengfa Fang
ICC1
2013 Performance bounds of compressed sensing recovery algorithms for sparse noisy signals
abstract
Recently, the performance bounds of the compressed sensing (CS) recovery algorithms have been investigated in the noisy setting. However, most of the papers only focus on the noisy measurement model where the signal is noiseless and the noise enters after the CS operation. The noisy signal model where both the signal and the compressed measurements are contaminated by the different noises is not considered. This paper works on the noisy signal model and provides the performance bounds for the following popular recovery algorithms: thresholding and orthogonal matching pursuit (OMP), Dantzig selector (DS) and basis pursuit denoising (BPDN). The performance of the recovery algorithms is quantified as the ℓ2distance between the reconstructed signal and the true noisy signal. Next, the impacts of the noise are analyzed on the basis of the quantified performance. The analysis results show that the effective way to restrain the impact of the noise is to choose the measurement matrix with low correlation between the columns or the rows. Finally, the theoretical bounds are verified with numerical simulations by calculating the mean-squared-error for the different noise variances. The simulation results show that OMP owns the better performance than the other three recovery algorithms under the noisy signal model.
Xiangling Li, Qimei Cui, Xiaofeng Tao 0001, Xianjun Yang, Waheed ur Rehman, Y. Jay Guo
WCNC4
2013 Energy-Efficient Distributed Data Storage for Wireless Sensor Networks Based on Compressed Sensing and Network Coding
abstract
Recently, distributed data storage (DDS) for Wireless Sensor Networks (WSNs) has attracted great attention, especially in catastrophic scenarios. Since power consumption is one of the most critical factors that affect the lifetime of WSNs, the energy efficiency of DDS in WSNs is investigated in this paper. Based on Compressed Sensing (CS) and network coding theories, we propose a Compressed Network Coding based Distributed data Storage (CNCDS) scheme by exploiting the correlation of sensor readings. The CNCDS scheme achieves high energy efficiency by reducing the total number of transmissions Nttot and receptions Nrtot during the data dissemination process. Theoretical analysis proves that the CNCDS scheme guarantees good CS recovery performance. In order to theoretically verify the efficiency of the CNCDS scheme, the expressions for Nttotand Nrtotare derived based on random geometric graphs (RGG) theory. Furthermore, based on the derived expressions, an adaptive CNCDS scheme is proposed to further reduce Nttot and Nrtot. Simulation results validate that, compared with the conventional ICStorage scheme, the proposed CNCDS scheme reduces Nttot, Nrtot, and the CS recovery mean squared error (MSE) by up to 55%, 74%, and 76% respectively. In addition, compared with the CNCDS scheme, the adaptive CNCDS scheme further reduces Nttotand Nrtotby up to 63% and 32% respectively.
Xianjun Yang, Xiaofeng Tao 0001, Eryk Dutkiewicz, Xiaojing Huang 0001, Y. Jay Guo, Qimei Cui
IEEE Trans. Wirel. Commun.1
2012 Random circulant orthogonal matrix based Analog Compressed Sensing
abstract
Analog Compressed Sensing (CS) has attracted considerable research interest in sampling area. One of the promising analog CS technique is the recently proposed Modulated Wideband Converter (MWC). However, MWC has a very high hardware complexity due to its parallel structure. To reduce the hardware complexity of MWC, this paper proposes a novel Random Circulant Orthogonal Matrix based Analog Compressed Sensing (RCOM-ACS) scheme. By circularly shifting the periodic mixing function, the RCOM-ACS scheme reduces the number of physical parallel channels from m to 1 at the cost of longer processing time, where m is in the order of several dozen to several hundred in MWC. It is proved that the m×M measurement matrix of RCOM-ACS scheme satisfies the Restricted Isometry Property (RIP) condition with probability 1-M-O(1)when m = O(rlog2Mlog3r), where M is the length of the periodic mixing function, r denotes the sparsity of the input signal. Furthermore, to make a good tradeoff between processing time and hardware complexity, a short processing time RCOM-ACS scheme is proposed in this paper. Simulation results show that, the proposed schemes outperform MWC in terms of recovery performance.
Xianjun Yang, Y. Jay Guo, Qimei Cui, Xiaofeng Tao 0001, Xiaojing Huang 0001
GLOBECOM1
2011 A Multistep Detection Scheme Based on Iteration for Cooperative Spectrum Sensing in Cognitive Radio
abstract
In cognitive radio (CR) networks, two main challenges faced by cooperative spectrum sensing are low detection performance and long detection time in low SNR scenario. To the best of our knowledge, little study analyzes these two problems simultaneously and thoroughly. Therefore, this paper proposes a novel spectrum sensing scheme, termed multistep detection (MD) scheme, which aims to resolve the above two problems. For MD scheme which consists of five steps, we focus on the following two key parts: two-threshold detection and iteration detection. For the two-threshold detection, we make theoretical analysis about the establishment of the two thresholds. And for the iteration detection, a novel spectrum sensing algorithm is proposed to improve the detection performance in low SNR. The iteration algorithm utilizes the following property to improve the signal's SNR: the variance of the mean of independent and identity Gaussian random variables is 1/K of one variable's variance. The simulation results indicate that the proposed MD scheme outperforms the traditional cooperative scheme both on the detection performance and the detection time significantly.
Xuefei Zhang 0003, Qimei Cui, Xianjun Yang, Xiaofeng Tao 0001
VTC Fall3
2011 Multi-antenna compressed wideband spectrum sensing for cognitive radio
abstract
This paper proposes a novel wideband spectrum sensing (WSS) scheme, termed multi-antenna compressed wideband spectrum sensing (MCWSS) scheme, which utilizes compressed sensing (CS) to reduce the extremely high sampling rate of wideband signal. Although there are studies on compressed wideband spectrum sensing, they only focus on single antenna signal. Since multi-antenna technology can enhance the detection performance, this paper investigates the multi-antenna scenario. However, existing CS recovery algorithms are designed only for single antenna signal and are not suitable for recovering multi-antenna signals. Therefore, the paper proposes two novel CS recovery algorithms from different angles, namely CRL2(combining relevance via L2norm) algorithm and CBS (combining before sampling) algorithm. The CRL2algorithm jointly recovers the multi-antenna signals and performs better than single antenna scenario. Whereas, CBS algorithm can significantly improves the recovery performance with an additional analog combining operation. Since existing WSS algorithms are too complicated, we devise a novel WSS algorithm, i.e. DA (divided-averaged) algorithm, which has good performance with low complexity. Simulation results show that the MCWSS scheme performs well at low sampling rate.
Xianjun Yang, Qimei Cui, Xiaofeng Tao 0001
WCNC1