VLDB 2026 Research / reviewers in the wild / expert
Shikun Zhang
dblp:83/3715
· DBLP profile ↗
125ranked-venue papers
7as first author
91since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 73 · 1 first-author · 57 since 2021Graphics, computer vision, multimedia, augmented reality and games · 34 · 3 first-author · 31 since 2021Software engineering, systems software and programming languages · 21 · 9 since 2021Databases, data management, data science and information retrieval · 11 · 9 since 2021Security and privacy · 8 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rethinking the Sampling Criteria in Reinforcement Learning for LLM Reasoning: A Competence-Difficulty Alignment PerspectiveabstractThe low sampling efficiency during the rollout phase poses a significant challenge to scaling reinforcement learning for large language model reasoning. Existing methods attempt to improve efficiency by scheduling problems based on problem difficulties. However, these approaches suffer from unstable and biased estimations of problem difficulty and fail to capture the alignment between model competence and problem difficulty in RL training, leading to suboptimal results. To address these challenges, we introduce Competence-Difficulty Alignment Sampling (CDAS). This approach allows for accurate and stable estimation of problem difficulties by aggregating historical performance discrepancies across problems. Subsequently, model competence is quantified to adaptively select problems whose difficulties align with the model's current competence using a fixed-point system. Extensive experiments in mathematical RL training show that CDAS consistently outperforms strong baselines, achieving the highest average accuracy of 45.89%. Furthermore, CDAS reduces the training step time overhead by 57.06% compared to the widely-used Dynamic Sampling strategy, verifying the efficiency of CDAS. Additional experiments on different tasks, model architectures, and model sizes demonstrate the generalization capability of CDAS. Deyang Kong, Xiangyu Xi, Wei Wang 0225, Jingang Wang, Shikun Zhang, Wei Ye 0004 |
AAAI | 7 |
| 2026 | Modeling Uncertainty Trends for Timely Retrieval in Dynamic RAGabstractDynamic retrieval-augmented generation (RAG) allows large language models (LLMs) to fetch external knowledge on demand, offering greater adaptability than static RAG. A central challenge in this setting lies in determining the optimal timing for retrieval. Existing methods often trigger retrieval based on low token-level confidence, which may lead to delayed intervention after errors have already propagated. We introduce Entropy-Trend Constraint (ETC), a training-free method that determines optimal retrieval timing by modeling the dynamics of token-level uncertainty. Specifically, ETC utilizes first- and second-order differences of the entropy sequence to detect emerging uncertainty trends, enabling earlier and more precise retrieval. Experiments on six QA benchmarks with three LLM backbones demonstrate that ETC consistently outperforms strong baselines while reducing retrieval frequency. ETC is particularly effective in domain-specific scenarios, exhibiting robust generalization capabilities. Ablation studies and qualitative analyses further confirm that trend-aware uncertainty modeling yields more effective retrieval timing. The method is plug-and-play, model-agnostic, and readily integrable into existing decoding pipelines. Implementation code is included in the supplementary materials. Bo Li 0099, Zhenghua Xu 0001, Shikun Zhang, Wei Ye 0004 |
AAAI | 5 |
| 2026 | L2-LoRA: Improving Low-Rank Adaptation with Layer-Specific RegularizationabstractFine-tuning large language models (LLMs) in a parameter-efficient manner while preserving their pre-trained world knowledge remains a significant challenge. While Low-Rank Adaptation (LoRA) and its variants effectively mitigate catastrophic forgetting, they do not fully eliminate the loss of critical pre-trained knowledge. In this work, we first analyze the layer-wise distribution of domain-specific knowledge within LLMs through knowledge localization, and empirically identify a clear layer-specific pattern: pre-trained world knowledge predominantly resides in lower layers, whereas knowledge relevant to downstream tasks is more concentrated in higher layers. Motivated by this observation, we propose L2-LoRA, a simple yet effective variant of LoRA that applies layer-specific L2 regularization to the LoRA weights during fine-tuning. Specifically, L2-LoRA imposes stronger regularization on lower layers to preserve pre-trained world knowledge, while allowing greater adaptation in higher layers to better align with downstream tasks. Experiments across multiple benchmarks show that L2-LoRA not only consistently outperforms vanilla LoRA in downstream performance, but also effectively mitigates catastrophic forgetting by retaining more pre-trained knowledge. Rui Xie 0003, Shikun Zhang |
AAAI | 3 |
| 2026 | ASKD: Reinforcement Learning-Style Knowledge Distillation with Quality-Adaptive SkewnessabstractKnowledge distillation (KD) is a widely adopted technique for transferring the capabilities of large teacher models to smaller student models, thereby significantly reducing inference costs and memory consumption. However, existing KD methods are all constrained by an inherent greedy optimization objective, rooted in the assumption of teacher superiority: "Trust all teacher-generated outputs (TGOs)" and "Distrust any student-generated outputs (SGOs) unsupported by the teacher". We propose ASKD, a novel KD method with adaptive skewness determined by sample quality, refining this objective to: "Learn TGOs proportionally to their quality, and distrust only low-quality unsupported SGOs". ASKD comprises three key components: (1) A reinforcement learning-style optimization formulation to mitigate the inherent approximation bias in sample-based Kullback-Leibler (KL) divergence approximations used by previous KD methods; (2) Well-designed quality supervision signals to map and achieve adaptive skewness in skewed KL loss, pioneering the usage of sample quality to adjust learning magnitudes; (3) A gradient-clip function on high-quality SGOs for findings that high-quality SGOs in KL loss fail to yield positive updates and even cause adverse effects on some samples. Extensive experiments indicate that ASKD builds high-performance student models across various tasks, including instruction following, mathematical reasoning, and code generation, outperforming state-of-the-art methods comprehensively and surpassing GRPO-like approaches that use advantages as multiplicative factors. We also provide detailed mathematical proofs demonstrating properties such as Lipschitz continuity of the update coefficient and uniform convergence of the loss function, ensuring theoretical rigor for key components of ASKD. Xiaoling Zhou, Yiyu Liu, Shikun Zhang, Wei Ye 0004 |
AAAI | 5 |
| 2026 | Retrieval as Generation: A Unified Framework with Self-Triggered Information PlanningabstractWe revisit retrieval-augmented generation (RAG) by embedding retrieval control directly into generation.Instead of treating retrieval as an external intervention, we express retrieval decisions within token-level decoding, enabling end-to-end coordination without additional controllers or classifiers.Under the paradigm of Retrieval as Generation, we propose GRIP (Generation-guided Retrieval with Information Planning), a unified framework in which the model regulates retrieval behavior through control-token emission.Central to GRIP is Self-Triggered Information Planning, which allows the model to decide when to retrieve, how to reformulate queries, and when to terminate, all within a single autoregressive trajectory.This design tightly couples retrieval and reasoning and supports dynamic multi-step inference with on-the-fly evidence integration.To supervise these behaviors, we construct a structured training set covering answerable, partially answerable, and multi-hop queries, each aligned with specific token patterns.Experiments on five QA benchmarks show that GRIP surpasses strong RAG baselines and is competitive with GPT-4o while using substantially fewer parameters. Bo Li 0099, Gexiang Fang, Shikun Zhang, Wei Ye 0004 |
ACL (1) | 4 |
| 2026 | Instruction Data Selection via Answer DivergenceabstractInstruction tuning relies on large instruction-response corpora whose quality and composition strongly affect downstream performance.We propose Answer Divergence-Guided Selection (ADG), which selects instruction data based on the geometric structure of multi-sample outputs.ADG draws several high-temperature generations per instruction, maps responses into an embedding space, and computes an output divergence score that jointly encodes dispersion magnitude and shape anisotropy.High scores correspond to instructions whose answers are both far apart and multi-modal, rather than clustered paraphrases along a single direction.Across two backbones and three public instruction pools, finetuning on only 10K ADG-selected examples consistently outperforms strong selectors on six benchmarks spanning reasoning, knowledge, and coding.Analyses further show that both dispersion magnitude and shape anisotropy are necessary, supporting answer divergence as a practical signal for instruction data selection.Code and appendix are included in the supplementary materials. Bo Li 0099, Shikun Zhang, Wei Ye 0004 |
ACL (1) | 3 |
| 2026 | LeLoRA: Learnable Low-Rank Adaptation of Large Language ModelsabstractFine-tuning large language models (LLMs) is an effective approach to enhancing their performance on specialized downstream tasks.Among the various techniques, low-rank adaptation has garnered significant attention due to its ability to maintain the full performance of fine-tuning while enhancing computational efficiency.However, existing approaches often rely on manually specified and fixed hyperparameters to identify the trainable components within weight matrices, resulting in suboptimal performance and low parameter efficiency.This paper presents a novel Learnable Low-Rank Adaptation (LeLoRA) framework that utilizes dynamically learned fine-tuning strategies to facilitate the effective adaptation of LLMs.Our framework integrates an LLM with a policy network that automatically and adaptively generates matrix-specific adaptation strategies to identify the trainable components of each weight matrix, taking into account their unique characteristics, such as singular values and matrix norms.A reinforcement learningbased optimization algorithm is then employed to iteratively update the LLM and the policy network, ensuring that the generated strategies adapt in real time to the evolving states of the LLM.Extensive experiments have been conducted across various natural language processing tasks.The results across ten different LLMs, ranging from 125M to 70B parameters, provide compelling evidence that LeLoRA consistently outperforms existing baselines in adapting LLMs. Xiaoling Zhou, Zhemg Lee, Wei Ye 0004, Shikun Zhang |
ACL (1) | 5 |
| 2026 | An Empirical Study on Influence-Based Pretraining Data Selection for Code Large Language Models
Chengli Xing, Zhengran Zeng, Gexiang Fang, Rui Xie 0003, Wei Ye 0004, Shikun Zhang |
ICPC | 6 |
| 2026 | Dynamic Knowledge Transfer for Mitigating Spurious Correlations in Deep Learning
Xiaoling Zhou, Zhemg Lee, Wei Ye 0004, Shikun Zhang |
Int. J. Comput. Vis. | 4 |
| 2026 | Label space compression and prompt alignment for collaborative inference with small and large language models
Shikun Zhang |
Neurocomputing | 3 |
| 2025 | Can You Really Trust Code Copilot? Evaluating Large Language Models from a Code Security PerspectiveabstractCode security and usability are both essential for various coding assistant applications driven by large language models (LLMs). Current code security benchmarks focus solely on single evaluation task and paradigm, such as code completion and generation, lacking comprehensive assessment across dimensions like secure code generation, vulnerability repair and discrimination. In this paper, we first propose CoV-Eval, a multi-task benchmark covering various tasks such as code completion, vulnerability repair, vulnerability detection and classification, for comprehensive evaluation of LLM code security. Besides, we developed VC-Judge, an improved judgment model that aligns closely with human experts and can review LLM-generated programs for vulnerabilities in a more efficient and reliable way. We conduct a comprehensive evaluation of 20 proprietary and open-source LLMs. Overall, while most LLMs identify vulnerable codes well, they still tend to generate insecure codes and struggle with recognizing specific vulnerability types and performing repairs. Extensive experiments and qualitative analyses reveal key challenges and optimization directions, offering insights for future research in LLM code security. Yutao Mou, Shikun Zhang, Wei Ye 0004 |
ACL (1) | 4 |
| 2025 | SymDPO: Boosting In-Context Learning of Large Multimodal Models with Symbol Demonstration Direct Preference OptimizationabstractAs language models continue to scale, Large Language Models (LLMs) have exhibited emerging capabilities in In-Context Learning (ICL), enabling them to solve language tasks by prefixing a few in-context demonstrations (ICDs) as context. Inspired by these advancements, researchers have extended these techniques to develop Large Multimodal Models (LMMs) with ICL capabilities. However, existing LMMs face a critical issue: they often fail to effectively leverage the visual context in multimodal demonstrations and instead simply follow textual patterns. This indicates that LMMs do not achieve effective alignment between multimodal demonstrations and model outputs. To address this problem, we propose Symbol Demonstration Direct Preference Optimization (SymDPO). Specifically, SymDPO aims to break the traditional paradigm of constructing multimodal demonstrations by using random symbols to replace text answers within instances. This forces the model to carefully understand the demonstration images and establish a relationship between the images and the symbols to answer questions correctly. We validate the effectiveness of this method on multiple benchmarks, demonstrating that with SymDPO, LMMs can more effectively understand the multimodal context within examples and utilize this knowledge to answer questions better. Code is available at https://github.com/APiaoG/SymDPO. Hongrui Jia, Chaoya Jiang, Haiyang Xu 0001, Wei Ye 0004, Mengfan Dong, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Shikun Zhang |
CVPR | 9 |
| 2025 | All-Optical Nonlinear Diffractive Deep Network for Ultrafast Image DenoisingabstractImage denoising poses a significant challenge in image processing, aiming to remove noise and artifacts from input images. However, current denoising algorithms implemented on electronic chips frequently encounter latency issues and demand substantial computational resources. In this paper, we introduce an all-optical Nonlinear Diffractive Denoising Deep Network (N3DNet) for image denoising at the speed of light. Initially, we incorporate an image encoding and pre-denoising module into the Diffractive Deep Neural Network and integrate a nonlinear activation function, termed the phase exponential linear function, after each diffractive layer, thereby boosting the network’s nonlinear modeling and denoising capabilities. Subsequently, we devise a new reinforcement learning algorithm called regularization-assisted deep Q-network to optimize N3DNet. Finally, leveraging 3D printing techniques, we fabricate N3DNet using the trained parameters and construct a physical experimental system for real-world applications. A new benchmark dataset, termed MIDD, is constructed for mode image denoising, comprising 120K pairs of noisy/noise-free images captured from real fiber communication systems across various transmission lengths. Through extensive simulation and real experiments, we validate that N3DNet outperforms both traditional and deep learning-based denoising approaches across various datasets. Remarkably, its processing speed is nearly 3,800 times faster than electronic chip-based methods. Xiaoling Zhou, Zhemg Lee, Wei Ye 0004, Rui Xie 0003, Guanju Peng, Shikun Zhang |
CVPR | 8 |
| 2025 | HaDeMiF: Hallucination Detection and Mitigation in Large Language ModelsabstractThe phenomenon of knowledge hallucinations has raised substantial concerns about the security and reliability of deployed large language models (LLMs). Current methods for detecting hallucinations primarily depend on manually designed individual metrics, such as prediction uncertainty and consistency, and fall short in effectively calibrating model predictions, thus constraining their detection accuracy and applicability in practical applications. In response, we propose an advanced framework, termed HaDeMiF, for detecting and mitigating hallucinations in LLMs. Specifically, hallucinations within the output and semantic spaces of LLMs are comprehensively captured through two compact networks—a novel, interpretable tree model known as the Deep Dynamic Decision Tree (D3T) and a Multilayer Perceptron (MLP)—which take as input a set of prediction characteristics and the hidden states of tokens, respectively. The predictions of LLMs are subsequently calibrated using the outputs from the D3T and MLP networks, aiming to mitigate hallucinations and enhance model calibration. HaDeMiF can be applied during both the inference and fine-tuning phases of LLMs, introducing less than 2% of the parameters relative to the LLMs through the training of two small-scale networks. Extensive experiments conclusively demonstrate the effectiveness of our framework in hallucination detection and model calibration across text generation tasks with responses of varying lengths. Xiaoling Zhou, Zhemg Lee, Wei Ye 0004, Shikun Zhang |
ICLR | 5 |
| 2025 | Reasoning Through Execution: Unifying Process and Outcome Rewards for Code GenerationabstractLarge Language Models excel at code generation yet struggle with complex programming tasks that demand sophisticated reasoning. To bridge this gap, traditional process supervision relies on learned reward models requiring costly training data and suffering from reward misalignment, while outcome supervision fails for complex tasks needing coordinated intermediate steps. We introduce Outcome Refining Process Supervision, which unifies process and outcome supervision by leveraging executable verification: a tree-structured search framework generates strategic alternatives, profiles execution metrics, and scores candidates via self-critique mechanisms that integrate runtime feedback with reasoning. Experiments across 5 models and 3 benchmarks show consistent gains, with 26.9% higher correctness and 42.2% improved code efficiency. The results demonstrate that ORPS enables LLMs to overcome local optima in code generation, suggesting a promising direction for combining verifiable outcomes with structured reasoning to tackle complex challenges. Zhuohao Yu 0001, Weizheng Gu, Yidong Wang 0003, Xingru Jiang, Zhengran Zeng, Jindong Wang 0001, Wei Ye 0004, Shikun Zhang |
ICML | 8 |
| 2025 | GETMusic: Generating Music Tracks with a Unified Representation and Diffusion FrameworkabstractSymbolic music generation aims to create musical notes, which can help users compose music, such as generating target instrument tracks based on provided source tracks. In practical scenarios where there’s a predefined ensemble of tracks and various composition needs, an efficient and effective generative model that can generate any target tracks based on the other tracks becomes crucial. However, previous efforts have fallen short in addressing this necessity due to limitations in their music representations and models. In this paper, we introduce a framework known as GETMusic, with ``GET'' standing for ``GEnerate music Tracks.'' This framework encompasses a novel music representation ``GETScore'' and a diffusion model ``GETDiff.'' GETScore represents musical notes as tokens and organizes tokens in a 2D structure, with tracks stacked vertically and progressing horizontally over time. At a training step, each track of a music piece is randomly selected as either the target or source. The training involves two processes: In the forward process, target tracks are corrupted by masking their tokens, while source tracks remain as the ground truth; in the denoising process, GETDiff is trained to predict the masked target tokens conditioning on the source tracks. Our proposed representation, coupled with the non-autoregressive generative model, empowers GETMusic to generate music with any arbitrary source-target track combinations.Our experiments demonstrate that the versatile GETMusic outperforms prior works proposed for certain specific composition tasks. Ang Lv, Xu Tan 0003, Peiling Lu, Wei Ye 0004, Shikun Zhang, Jiang Bian 0002, Rui Yan 0001 |
IJCAI | 5 |
| 2025 | Robustness to Spurious Correlations via Dynamic Knowledge TransferabstractSpurious correlations pose a significant challenge to the robustness of statistical models, often resulting in unsatisfactory performance when distributional shifts occur between training and testing data. To address this, we propose to transfer knowledge across spuriously correlated categories within the deep feature space. Specifically, samples' deep features are enriched using semantic vectors extracted from both their respective category distributions and those of their spuriously correlated counterparts, enabling the generation of diverse class-specific factual and counterfactual augmented deep features. We then demonstrate the feasibility of optimizing a surrogate robust loss instead of conducting explicit augmentations by considering an infinite number of augmentations. As spurious correlations between samples and classes evolve during training, we develop a reinforcement learning-based training framework called Dynamic Knowledge Transfer (DKT) to facilitate dynamic adjustments in the direction and intensity of knowledge transfer. Within this framework, a target network is trained using the derived robust loss to enhance robustness, while a strategy network generates sample-wise augmentation strategies in a dynamic and automatic way. Extensive experiments validate the effectiveness of the DKT framework in mitigating spurious correlations, achieving state-of-the-art performance across three typical learning scenarios susceptible to such correlations. Xiaoling Zhou, Wei Ye 0004, Zhemg Lee, Shikun Zhang |
IJCAI | 4 |
| 2025 | R3-Bench: Reproducible Real-world Reverse Engineering Dataset for Symbol RecoveryabstractSymbol recovery in reverse engineering is crucial for restoring variable and data structure information in compiled binaries. While learning-based methods have shown promise in recovering both semantic information (names and types) and syntactic information (shapes), they require comprehensive datasets where expressions in binary code are precisely aligned with their source code equivalents. Current techniques for generating such alignments struggle with complex data access patterns, resulting in incomplete training data and consequently hampering model performance and recovery accuracy. We present AST-Align, a novel technique unifying alignment of variables and struct access expressions across multiple architectures (x86 and ARM) and languages (C/C++/Rust). AST-Align significantly improves the number of generated ground truths, capturing four times more struct fields than previous methods. Using this algorithm, we develop R3-Bench, a metadata-rich, extensible dataset with explicit project inclusion criteria and reproducible processing pipeline, comprising over 10 million functions across multiple architectures. Our evaluation establishes baseline performance by testing various approaches from n-gram models to Large Language Models. The results show that while general LLMs initially perform poorly, their effectiveness dramatically improves with proper demonstration. R3-Bench provides a robust foundation for assessing model capabilities and serves as a valuable reference for future symbol recovery research. Muzhi Yu, Zhengran Zeng, Wei Ye 0004, Jinan Sun, Xiaolong Bai, Shikun Zhang |
ASE | 6 |
| 2025 | VLM-R³: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-ThoughtabstractRecently, reasoning-based MLLMs have achieved a degree of success in generating long-form textual reasoning chains. However, they still struggle with complex tasks that necessitate dynamic and iterative focusing on and revisiting of visual regions to achieve precise grounding of textual reasoning in visual evidence. We introduce VLM-R³ (Visual Language Model with Region Recognition, Reasoning, and Refinement ), a framework that equips an MLLM with the ability to (i) decide when additional visual evidence is needed, (ii) determine where to ground within the image, and (iii) seamlessly weave the relevant sub-image content back into an interleaved chain-of-thought. The core of our method is \textbf{Region-Conditioned Reinforcement Policy Optimization (R-GRPO)}, a training paradigm that rewards the model for selecting informative regions, formulating appropriate transformations (e.g. crop, zoom), and integrating the resulting visual context into subsequent reasoning steps. To bootstrap this policy, we compile a modest but carefully curated Visuo-Lingual Interleaved Rationale (VLIR) corpus that provides step-level supervision on region selection and textual justification. Extensive experiments on MathVista, ScienceQA, and other benchmarks show that VLM-R$^3$ sets a new state of the art in zero-shot and few-shot settings, with the largest gains appearing on questions demanding subtle spatial reasoning or fine-grained visual cue extraction. Chaoya Jiang, Yongrui Heng, Wei Ye 0004, Haiyang Xu 0001, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Shikun Zhang |
NeurIPS | 8 |
| 2025 | SAEMark: Steering Personalized Multilingual LLM Watermarks with Sparse AutoencodersabstractWatermarking LLM-generated text is critical for content attribution and misinformation prevention, yet existing methods compromise text quality and require white-box model access with logit manipulation or training, which exclude API-based models and multilingual scenarios. We propose SAEMark, an **inference-time framework** for *multi-bit* watermarking that embeds personalized information through *feature-based rejection sampling*, fundamentally different from logit-based or rewriting-based approaches: we **do not modify model outputs directly** and require only **black-box access**, while naturally supporting multi-bit message embedding and generalizing across diverse languages and domains. We instantiate the framework using *Sparse Autoencoders* as deterministic feature extractors and provide theoretical worst-case analysis relating watermark accuracy to computational budget. Experiments across 4 datasets demonstrate strong watermarking performance on English, Chinese, and code while preserving text quality. SAEMark establishes a new paradigm for **scalable, quality-preserving watermarks** that work seamlessly with closed-source LLMs across languages and domains. Zhuohao Yu 0001, Xingru Jiang, Weizheng Gu, Yidong Wang 0003, Qingsong Wen, Shikun Zhang, Wei Ye 0004 |
NeurIPS | 6 |
| 2025 | Boosting Resilience of Large Language Models through Causality-Driven Robust OptimizationabstractLarge language models (LLMs) have achieved remarkable achievements across diverse applications; however, they remain plagued by spurious correlations and the generation of hallucinated content. Despite extensive efforts to enhance the resilience of LLMs, existing approaches either rely on indiscriminate fine-tuning of all parameters, resulting in parameter inefficiency and lack of specificity, or depend on post-processing techniques that offer limited adaptability and flexibility. This study introduces a novel Causality-driven Robust Optimization (CdRO) approach that selectively updates model components sensitive to causal reasoning, enhancing model causality while preserving valuable pretrained knowledge to mitigate overfitting. Our method begins by identifying the parameter components within LLMs that capture causal relationships, achieved through comparing the training dynamics of parameter matrices associated with the original samples, as well as augmented counterfactual and paraphrased variants. These comparisons are then fed into a lightweight logistic regression model, optimized in real time to dynamically identify and adapt the causal components within LLMs. The identified parameters are subsequently optimized using an enhanced policy optimization algorithm, where the reward function is designed to jointly promote both model generalization and robustness. Extensive experiments across various tasks using twelve different LLMs demonstrate the superior performance of our framework, underscoring its significant effectiveness in reducing the model’s dependence on spurious associations and mitigating hallucinations. Xiaoling Zhou, Zhemg Lee, Yuncheng Hua, Chengli Xing, Wei Ye 0004, Flora D. Salim, Shikun Zhang |
NeurIPS | 8 |
| 2025 | Click Without Compromise: Online Advertising Measurement via Per User Differential PrivacyabstractOnline advertising is a cornerstone of the Internet ecosystem, with advertising measurement playing a crucial role in optimizing efficiency. Ad measurement entails attributing desired behaviors, such as purchases, to ad exposures across various platforms, necessitating the collection of user activities across these platforms. As this practice faces increasing restrictions due to rising privacy concerns, safeguarding user privacy in this context is imperative. Our work is the first to formulate the real-world challenge of advertising measurement systems with real-time reporting of streaming data in advertising campaigns. We introduce AdsBPC, a novel user-level differential privacy protection scheme for online advertising measurement results. This approach optimizes global noise power and results in a non-identically distributed noise distribution that preserves differential privacy while enhancing measurement accuracy. Through experiments on both real-world advertising campaigns and synthetic datasets, AdsBPC achieves a 33% to 95% increase in accuracy over existing streaming DP mechanisms applied to advertising measurement. This highlights our method's effectiveness in achieving superior accuracy alongside a formal privacy guarantee, thereby advancing the state-of-the-art in privacy-preserving advertising measurement. Yingtai Xiao, Shikun Zhang, Wanrong Zhang 0004, Danfeng Zhang, Daniel Kifer |
SP | 3 |
| 2025 | Mitigating spurious correlations with causal logit perturbation
Xiaoling Zhou, Wei Ye 0004, Rui Xie 0003, Shikun Zhang |
Inf. Sci. | 4 |
| 2025 | 3D surface reconstruction with enhanced high-frequency details
Shikun Zhang, Yiqun Wang 0001, Cunjian Chen, Yong Li 0023, Qiuhong Ke |
J. Vis. Commun. Image Represent. | 1 |
| 2025 | Valuing Training Data via Causal Inference for In-Context LearningabstractIn-context learning (ICL) empowers large pre-trained language models (PLMs) to predict outcomes for unseen inputs without parameter updates. However, the efficacy of ICL heavily relies on the choice of demonstration examples. Randomly selecting from the training set frequently leads to inconsistent performance. Addressing this challenge, this study takes a novel approach by focusing on training data valuation through causal inference. Specifically, we introduce the concept of average marginal effect (AME) to quantify the contribution of individual training samples to ICL performance, encompassing both its generalization and robustness. Drawing inspiration from multiple treatment effects and randomized experiments, we initially sample diverse training subsets to construct prompts and evaluate the ICL performance based on these prompts. Subsequently, we employ Elastic Net regression to collectively estimate the AME values for all training data, considering subset compositions and inference performance. Ultimately, we prioritize samples with the highest values to prompt the inference of the test data. Across various tasks and with seven PLMs ranging in size from 0.8B to 33B, our approach consistently achieves state-of-the-art performance. Particularly, it outperforms Vanilla ICL and the best-performing baseline by an average of 14.1% and 5.2%, respectively. Moreover, prioritizing the most valuable samples for prompting leads to a significant enhancement in performance stability and robustness across various learning scenarios. Impressively, the valuable samples exhibit transferability across diverse PLMs and generalize well to out-of-distribution tasks. Xiaoling Zhou, Wei Ye 0004, Zhemg Lee, Lei Zou 0001, Shikun Zhang |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2024 | TiMix: Text-Aware Image Mixing for Effective Vision-Language Pre-trainingabstractSelf-supervised Multi-modal Contrastive Learning (SMCL) remarkably advances modern Vision-Language Pre-training (VLP) models by aligning visual and linguistic modalities. Due to noises in web-harvested text-image pairs, however, scaling up training data volume in SMCL presents considerable obstacles in terms of computational cost and data inefficiency. To improve data efficiency in VLP, we propose Text-aware Image Mixing (TiMix), which integrates mix-based data augmentation techniques into SMCL, yielding significant performance improvements without significantly increasing computational overhead. We provide a theoretical analysis of TiMix from a mutual information (MI) perspective, showing that mixed data samples for cross-modal contrastive learning implicitly serve as a regularizer for the contrastive loss. The experimental results demonstrate that TiMix exhibits a comparable performance on downstream tasks, even with a reduced amount of training data and shorter training time, when benchmarked against existing methods. This work empirically and theoretically demonstrates the potential of data mixing for data-efficient and computationally viable VLP, benefiting broader VLP model adoption in practical scenarios. Our code is available on https://github.com/chaoyajiang/TiMiX/tree/main. Chaoya Jiang, Wei Ye 0004, Haiyang Xu 0001, Qinghao Ye, Ming Yan 0008, Ji Zhang 0011, Shikun Zhang |
AAAI | 7 |
| 2024 | Labels Need Prompts Too: Mask Matching for Natural Language Understanding TasksabstractTextual label names (descriptions) are typically semantically rich in many natural language understanding (NLU) tasks. In this paper, we incorporate the prompting methodology, which is widely used to enrich model input, into the label side for the first time. Specifically, we propose a Mask Matching method, which equips an input with a prompt and its label with another, and then makes predictions by matching their mask representations. We evaluate our method extensively on 8 NLU tasks with 14 datasets. The experimental results show that Mask Matching significantly outperforms its counterparts of fine-tuning and conventional prompt-tuning, setting up state-of-the-art performances in several datasets. Mask Matching is particularly good at handling NLU tasks with large label counts and informative label names. As pioneering efforts that investigate the label-side prompt, we also discuss open issues for future study. Bo Li 0099, Wei Ye 0004, Quansen Wang, Shikun Zhang |
AAAI | 5 |
| 2024 | KIEval: A Knowledge-grounded Interactive Evaluation Framework for Large Language ModelsabstractZhuohao Yu, Chang Gao, Wenjin Yao, Yidong Wang, Wei Ye, Jindong Wang, Xing Xie, Yue Zhang, Shikun Zhang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zhuohao Yu 0001, Wenjin Yao, Yidong Wang 0003, Wei Ye 0004, Jindong Wang 0001, Xing Xie 0001, Yue Zhang 0004, Shikun Zhang |
ACL (1) | 9 |
| 2024 | Enhancing In-Context Learning via Implicit Demonstration AugmentationabstractXiaoling Zhou, Wei Ye, Yidong Wang, Chaoya Jiang, Zhemg Lee, Rui Xie, Shikun Zhang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Xiaoling Zhou, Wei Ye 0004, Yidong Wang 0003, Chaoya Jiang, Zhemg Lee, Rui Xie 0003, Shikun Zhang |
ACL (1) | 7 |
| 2024 | A Student Emotions Recognition Based Online Texts with Large Language Models for Learning Predictionabstractwith the rapid development of artificial intelligence (AI) and Large Language Models (LLM) technologies, it offers innovative approaches to model and analyze to predict student performance with educational behavior and learning data. But the challenges of data diversity, technical complex, and lack of semantic comprehension ability that limited the use of AI-based tools for learning performance prediction. In the paper, student emotions recognition based online learning data especially forum texts using LLMs is proposed, and an improved online learning performance prediction with online learning data, including student emotions using Signed Graph Neural Networks (SGNN) and LLMs, namely ER-SGNN-LLM, is constructed. In the model, the keywords are extracted from both forum texts and answers generated by students using LLM, and the keywords extracted from student’s forum texts are used for student emotions recognition, the keywords extracted from student’s answers, as well as the student emotions are used for student learning prediction. The relationships between each keyword extracted from texts and knowledge points of the course are encoded with SGNN. A graphic contrastive learning model is used to handle the noise in the dataset caused by students' subjective reasons. The combination of GNN and LLM is used to student emotions recognition and learning prediction. To verify the performances of the model, many experiments are conducted using the public dataset. The results demonstrate that, the model achieved better effectiveness in learning prediction, compared with other models. The values of F1 score of the proposed model is improved 0.032 compared that of the exist SGNN-LLM model. Guangyu Fan, Shikun Zhang, Songlin Cheng, Dingyu Yang |
BIBM | 2 |
| 2024 | Hallucination Augmented Contrastive Learning for Multimodal Large Language ModelabstractMulti-modal large language models (MLLMs) have been shown to efficiently integrate natural language with visual information to handle multi-modal tasks. However, MLLMs still face a fundamental limitation of hallucinations, where they tend to generate erroneous or fabricated information. In this paper, we address hallucinations in MLLMs from a novel perspective of representation learning. We first analyzed the representation distribution of textual and visual tokens in MLLM, revealing two important findings: 1) there is a significant gap between textual and visual representations, indicating unsatisfactory cross-modal representation alignment; 2) representations of texts that contain and do not contain hallucinations are entangled, making it challenging to distinguish them. These two observations inspire us with a simple yet effective method to mitigate hallucinations. Specifically, we introduce contrastive learning into MLLMs and use text with hallucination as hard negative examples, naturally bringing representations of non-hallucinative text and visual samples closer while pushing way representations of non-hallucinating and hallucinative text. We evaluate our method quantitatively and qualitatively, showing its effectiveness in reducing hallucination occurrences and improving performance across multiple benchmarks. On the MMhal-Bench benchmark, our method obtains a 34.66% /29.5% improvement over the baseline MiniGPT-4/LLaVA. Our code is available on https://github.com/X-PLUG/mPLUG-HalOwl/tree/main/hacl. Chaoya Jiang, Haiyang Xu 0001, Mengfan Dong, Wei Ye 0004, Ming Yan 0008, Qinghao Ye, Ji Zhang 0011, Fei Huang 0002, Shikun Zhang |
CVPR | 10 |
| 2024 | PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning OptimizationabstractInstruction tuning large language models (LLMs) remains a challenging task, owing to the complexity of hyperparameter selection and the difficulty involved in evaluating the tuned models. To determine the optimal hyperparameters, an automatic, robust, and reliable evaluation benchmark is essential. However, establishing such a benchmark is not a trivial task due to the challenges associated with evaluation accuracy and privacy protection. In response to these challenges, we introduce a judge large language model, named PandaLM, which is trained to distinguish the superior model given several LLMs. PandaLM's focus extends beyond just the objective correctness of responses, which is the main focus of traditional evaluation datasets. It addresses vital subjective factors such as relative conciseness, clarity, adherence to instructions, comprehensiveness, and formality. To ensure the reliability of PandaLM, we collect a diverse human-annotated test dataset, where all contexts are generated by humans and labels are aligned with human preferences. Our findings reveal that PandaLM-7B offers a performance comparable to both GPT-3.5 and GPT-4. Impressively, PandaLM-70B surpasses their performance. PandaLM enables the evaluation of LLM to be fairer but with less cost, evidenced by significant improvements achieved by models tuned through PandaLM compared to their counterparts trained with default Alpaca's hyperparameters. In addition, PandaLM does not depend on API-based evaluations, thus avoiding potential data leakage. Yidong Wang 0003, Zhuohao Yu 0001, Wenjin Yao, Zhengran Zeng, Linyi Yang, Cunxiang Wang, Hao Chen 0102, Chaoya Jiang, Rui Xie 0003, Jindong Wang 0001, Xing Xie 0001, Wei Ye 0004, Shikun Zhang, Yue Zhang 0004 |
ICLR | 13 |
| 2024 | NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion ModelsabstractWhile recent large-scale text-to-speech (TTS) models have achieved significant progress, they still fall shorts in speech quality, similarity, and prosody. Considering that speech intricately encompasses various attributes (e.g., content, prosody, timbre, and acoustic details) that pose significant challenges for generation, a natural idea is to factorize speech into individual subspaces representing different attributes and generate them individually. Motivated by it, we propose a TTS system with novel factorized diffusion models to generate natural speech in a zero-shot way. Specifically, 1) we design a neural codec with factorized vector quantization (FVQ) to disentangle speech waveform into subspaces of content, prosody, timbre, and acoustic details; 2) we propose a factorized diffusion model, which generates attributes in each subspace following its corresponding prompt. With this factorization design, our method can effectively and efficiently model the intricate speech with disentangled subspaces in a divide-and-conquer way. Experimental results show that our method outperforms the state-of-the-art TTS systems on quality, similarity, prosody, and intelligibility. Zeqian Ju, Yuancheng Wang, Xu Tan 0003, Detai Xin, Dongchao Yang, Eric Liu 0006, Yichong Leng, Kaitao Song, Siliang Tang, Zhizheng Wu 0001, Tao Qin 0001, Xiang-Yang Li 0001, Wei Ye 0004, Shikun Zhang, Jiang Bian 0002, Lei He 0005, Jinyu Li 0001, Sheng Zhao 0002 |
ICML | 15 |
| 2024 | Improving Long-Tail Vulnerability Detection Through Data Augmentation Based on Large Language ModelsabstractThe ability of automatic vulnerability detection models largely depends on the dataset used for training. However, annotating these datasets is costly and time-consuming, leading to a scarcity of labeled samples, particularly for diverse CWE types. This scarcity results in a pronounced long-tail issue, where less common types are underrepresented, thus diminishing the model's effectiveness in detecting them. In this paper, we address this challenge by employing large language models' (LLMs) generative and reasoning capabilities to create the necessary training samples for detecting less common vulnerability types. Specifically, we use GPT-4 to generate targeted samples of these types. After a well-defined self-filtering process, these samples are incorporated into the training of detection models. Extensive experiments show that our approach significantly enhances vulnerability detection capabilities, especially for long-tail vulnerabilities, across a variety of detection models, including both traditional deep learning and modern LLM-based detection models. Further, our comparative tests on unseen projects indicate that models trained with our generated data can identify more real-world vulnerabilities than traditional methods, proving the practicality and generalizability of our approach in real settings. The code and the generated dataset are publicly available at https://github.com/LuckyDengXiao/LERT. Fuyao Duan, Rui Xie 0003, Wei Ye 0004, Shikun Zhang |
ICSME | 5 |
| 2024 | Boosting Model Resilience via Implicit Adversarial Data Augmentation
Xiaoling Zhou, Wei Ye 0004, Zhemg Lee, Rui Xie 0003, Shikun Zhang |
IJCAI | 5 |
| 2024 | CoderUJB: An Executable and Unified Java Benchmark for Practical Programming ScenariosabstractIn the evolving landscape of large language models (LLMs) tailored for software engineering, the need for benchmarks that accurately reflect real-world development scenarios is paramount. Current benchmarks are either too simplistic or fail to capture the multi-tasking nature of software development. To address this, we introduce CoderUJB, a new benchmark designed to evaluate LLMs across diverse Java programming tasks that are executable and reflective of actual development scenarios, acknowledging Java's prevalence in real-world software production. CoderUJB comprises 2,239 programming questions derived from 17 real open-source Java projects and spans five practical programming tasks. Our empirical study on this benchmark investigates the coding abilities of various open-source and closed-source LLMs, examining the effects of continued pre-training in specific programming languages code and instruction fine-tuning on their performance. The findings indicate that while LLMs exhibit strong potential, challenges remain, particularly in non-functional code generation (e.g., test generation and defect detection). Importantly, our results advise caution in the specific programming languages continued pre-training and instruction fine-tuning, as these techniques could hinder model performance on certain tasks, suggesting the need for more nuanced strategies. CoderUJB thus marks a significant step towards more realistic evaluations of programming capabilities in LLMs, and our study provides valuable insights for the future development of these models in software engineering. Zhengran Zeng, Yidong Wang 0003, Rui Xie 0003, Wei Ye 0004, Shikun Zhang |
ISSTA | 5 |
| 2024 | Hal-Eval: A Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language ModelsabstractLarge Vision-Language Models (LVLMs) exhibit remarkable capabilities but struggle with ''hallucinations''-inconsistencies between images and their descriptions. Previous hallucination evaluation studies on LVLMs have identified hallucinations in terms of objects, attributes, and relations but overlooked complex hallucinations that create an entire narrative around a fictional entity. In this paper, we introduce a refined taxonomy of hallucinations, featuring a new category: Event Hallucination. We then utilize advanced LLMs to generate and filter fine-grained hallucinatory data consisting of various types of hallucinations, with a particular focus on event hallucinations, laying the groundwork for integrating discriminative and generative evaluation methods within our universal evaluation framework. The proposed benchmark distinctively assesses LVLMs' ability to tackle a broad spectrum of hallucinations, making it a reliable and comprehensive tool for gauging LVLMs' efficacy in handling hallucinations. We will release our code and data. Chaoya Jiang, Hongrui Jia, Mengfan Dong, Wei Ye 0004, Haiyang Xu 0001, Ming Yan 0008, Ji Zhang 0011, Shikun Zhang |
ACM Multimedia | 8 |
| 2024 | MaVEn: An Effective Multi-granularity Hybrid Visual Encoding Framework for Multimodal Large Language ModelabstractThis paper presents MaVEn, an innovative Multi-granularity Visual Encoding framework designed to enhance the capabilities of Multimodal Large Language Models (MLLMs) in multi-image reasoning. Current MLLMs primarily focus on single-image visual understanding, limiting their ability to interpret and integrate information across multiple images. MaVEn addresses this limitation by combining discrete visual symbol sequences, which abstract coarse-grained semantic concepts, with traditional continuous representation sequences that model fine-grained features. This dual approach bridges the semantic gap between visual and textual data, thereby improving the model's ability to process and interpret information from multiple images effectively. Additionally, we design a dynamic reduction mechanism by for long-sequence continuous features to enhance multi-image processing efficiency. Experimental results demonstrate that MaVEn significantly enhances MLLMs' understanding in complex multi-image scenarios, while also improving performance in single-image contexts. Chaoya Jiang, Hongrui Jia, Haiyang Xu 0001, Wei Ye 0004, Mengfan Dong, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Shikun Zhang |
NeurIPS | 9 |
| 2024 | SG-Bench: Evaluating LLM Safety Generalization Across Diverse Tasks and Prompt TypesabstractEnsuring the safety of large language model (LLM) applications is essential for developing trustworthy artificial intelligence. Current LLM safety benchmarks have two limitations. First, they focus solely on either discriminative or generative evaluation paradigms while ignoring their interconnection. Second, they rely on standardized inputs, overlooking the effects of widespread prompting techniques, such as system prompts, few-shot demonstrations, and chain-of-thought prompting. To overcome these issues, we developed SG-Bench, a novel benchmark to assess the generalization of LLM safety across various tasks and prompt types. This benchmark integrates both generative and discriminative evaluation tasks and includes extended data to examine the impact of prompt engineering and jailbreak on LLM safety. Our assessment of 3 advanced proprietary LLMs and 10 open-source LLMs with the benchmark reveals that most LLMs perform worse on discriminative tasks than generative ones, and are highly susceptible to prompts, indicating poor generalization in safety alignment. We also explain these findings quantitatively and qualitatively to provide insights for future research. Yutao Mou, Shikun Zhang, Wei Ye 0004 |
NeurIPS | 2 |
| 2024 | AutoSurvey: Large Language Models Can Automatically Write SurveysabstractThis paper introduces AutoSurvey, a speedy and well-organized methodology for automating the creation of comprehensive literature surveys in rapidly evolving fields like artificial intelligence. Traditional survey paper creation faces challenges due to the vast volume and complexity of information, prompting the need for efficient survey methods. While large language models (LLMs) offer promise in automating this process, challenges such as context window limitations, parametric knowledge constraints, and the lack of evaluation benchmarks remain. AutoSurvey addresses these challenges through a systematic approach that involves initial retrieval and outline generation, subsection drafting by specialized LLMs, integration and refinement, and rigorous evaluation and iteration. Our contributions include a comprehensive solution to the survey problem, a reliable evaluation method, and experimental validation demonstrating AutoSurvey's effectiveness. Yidong Wang 0003, Wenjin Yao, Xin Zhang 0097, Zhen Wu 0002, Meishan Zhang, Xinyu Dai, Min Zhang 0005, Qingsong Wen, Wei Ye 0004, Shikun Zhang, Yue Zhang 0004 |
NeurIPS | 12 |
| 2024 | Understanding How to Inform Blind and Low-Vision Users about Data Privacy through Privacy Question Answering Assistants
Yuanyuan Feng, Abhilasha Ravichander, Yaxing Yao, Shikun Zhang, Rex Chen, Shomir Wilson, Norman M. Sadeh |
USENIX Security Symposium | 4 |
| 2024 | iHunter: Hunting Privacy Violations at Scale in the Software Supply Chain on iOS
Dexin Liu, Yue Xiao 0007, Chaoqi Zhang 0006, Kaitao Xie, Xiaolong Bai, Shikun Zhang, Luyi Xing |
USENIX Security Symposium | 6 |
| 2024 | ALANCA: Active Learning Guided Adversarial Attacks for Code Comprehension on Diverse Pre-trained and Large Language ModelsabstractNeural code models have demonstrated their efficacy across a range of code comprehension tasks, including vulnerability detection, code classification, automatic code summarization, completion, clone detection, etc. Yet, a substantial gap exists in our understanding of the robustness of models in the realm of code comprehension and its associated applications. To probe and illuminate the robustness of code, recent efforts have sought to employ NLP-like techniques to craft adversarial code instances, primarily by perturbing variable and token names. It's worth noting that the semantics of source code predominantly surface through its structural elements, such as abstract syntax trees and control flow graphs, which fundamentally differ from natural languages. The question remains open: Can we perturb the structural aspects of code while preserving its semantics, thereby generating more disruptive adversarial examples that elude current structural-unaware approaches? Moreover, orchestrating adaptive adversarial attacks on diverse neural code models with varying architectures poses formidable challenges, especially in real-world scenarios characterized by constraints on target model access and querying. In this paper, we introduce ALANCA, an active-learning guided adversarial attack framework tailored for neural code models. Leveraging semantic-preserving translations, combined with an adaptive adversarial discriminator and token selector, ALANCA excels in executing adversarial attacks with high success rates, exceptional generation quality, and adaptability across different target models. We substantiate ALANCA's efficacy through comprehensive evaluations across four distinct code comprehension tasks, demonstrating its ability to effectively confound a range of neural models, including pre-trained models and LLMs used in software engineering. Dexin Liu, Shikun Zhang |
SANER | 2 |
| 2024 | SWAT4J: Generating System Call Allowlist for Java Container Attack Surface ReductionabstractWith the widespread use of container technology, attackers may invade the kernel by maliciously executing certain system calls, causing damage to the host and other containers. In order to reduce the attack surface of the underlying system, Docker supports specifying a container's allowlist of system calls with seccomp configurations. Java is a mainstream programming language used by the container projects in the Docker Hub, but how to generate the allowlist of system calls for Java containers is still an open question. Firstly, most of previous efforts about container allowlist of system calls focused on the C/C++ binary code rather than Java bytecode. Secondly, some existing works on Java bytecode mainly paid attention to the security vulnerabilities analysis, and cannot be used to analyze system calls required by Java programs. In this paper, we propose the first bytecode-based system call analysis approach, named SWAT 4J, tailored for Java containers operating on x86_64 architecture. SWAT4J can generate the allowlist of system calls required for Java containers by combining static and dynamic analysis. For static analysis, SWAT4J can identify the indirect calling relationships between Java bytecode and system calls, and determine the system calls required for a containerized application. For dynamic analysis, SWAT4J can trace the system calls required for container startup. The seccomp configuration file is optimized through the combining set of system calls. In the end, we experimented with 5 types of popular open source Java containers projects from Docker Official Images. Compared to 323 system calls in Ubuntu 16.04, SWAT4J successfully reduce the number of system calls by 56.04%-59.44%, and reduce the probability of vulnerabilities without affecting the functionality of the container. Yijiang Xu, Muxian Zhou, Shikun Zhang, Zhonghai Wu |
SANER | 4 |
| 2024 | Exploring Vision-Language Models for Imbalanced Learning
Yidong Wang 0003, Zhuohao Yu 0001, Jindong Wang 0001, Qiang Heng, Hao Chen 0102, Wei Ye 0004, Rui Xie 0003, Xing Xie 0001, Shikun Zhang |
Int. J. Comput. Vis. | 9 |
| 2024 | MAC: Maximal Cliques for 3D RegistrationabstractThis paper presents a 3D registration method with maximal cliques (MAC) for 3D point cloud registration (PCR). The key insight is to loosen the previous maximum clique constraint and mine more local consensus information in a graph for accurate pose hypotheses generation: 1) A compatibility graph is constructed to render the affinity relationship between initial correspondences. 2) We search for maximal cliques in the graph, each representing a consensus set. 3) Transformation hypotheses are computed for the selected cliques by the SVD algorithm and the best hypothesis is used to perform registration. In addition, we present a variant of MAC if given overlap prior, called MAC-OP. Overlap prior further enhances MAC from many technical aspects, such as graph construction with re-weighted nodes, hypotheses generation from cliques with additional constraints, and hypothesis evaluation with overlap-aware weights. Extensive experiments demonstrate that both MAC and MAC-OP effectively increase registration recall, outperform various state-of-the-art methods, and boost the performance of deep-learned methods. For instance, MAC combined with GeoTransformer achieves a state-of-the-art registration recall of [Formula: see text] on 3DMatch / 3DLoMatch. We perform synthetic experiments on 3DMatch-LIR / 3DLoMatch-LIR, a dataset with extremely low inlier ratios for 3D registration in ultra-challenging cases. Jiaqi Yang 0002, Xiyu Zhang 0001, Peng Wang 0015, Yulan Guo, Kun Sun 0002, Qiao Wu, Shikun Zhang, Yanning Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2024 | Toward Meta-Shape-Based Multi-View 3D Point Cloud Registration: An EvaluationabstractReducing cumulative registration error is critical to accurate 3D multi-view registration. Meta-shape based methods optimize rigid transformations of point clouds by iteratively registering each point cloud with a meta-shape, which remain popular solutions to 3D multi-view registration. However, the merits and demerits of existing meta-shape based methods remain unclear. Moreover, we argue that simpler meta-shape based solutions can achieve even better performance. To this end, we evaluate seven representative meta-shape based methods in this work, including four existing ones and three modified ones, in order to investigate the problem of defining a good meta-shape. In particular, we first abstract the main steps of considered methods. Then, experiments on both object and scene datasets with real and synthetic cumulative registration errors are deployed for an in-depth evaluation. Finally, based on the experimental outcomes, we give a discussion on the advantages and limitations of meta-shape based methods. We demonstrate prior works have used unnecessarily complicated techniques for cumulative error elimination and our slightly modified simpler solutions can achieve competitive performance on experimental datasets. Shikun Zhang, Jiaqi Yang 0002, Zhaoshuai Qi, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | DIOR: Learning to Hash With Label Noise Via Dual Partition and Contrastive LearningabstractDue to the excellent computing efficiency, learning to hash has acquired broad popularity for Big Data retrieval. Although supervised hashing methods have achieved promising performance recently, they presume that all training samples are appropriately annotated. Unfortunately, label noise is ubiquitous owing to erroneous annotations in real-world applications, which could seriously deteriorate the retrieval performance due to imprecise supervised guidance and severe memorization of noisy data. Here we propose a comprehensive method DIOR to handle the difficulties of learning to hash with label noise. DIOR performs partitions from two complementary levels, namely sample level and parameter level. On the one hand, DIOR divides the dataset into a labeled set with clean samples and an unlabeled set with noisy samples using an ensemble of perturbed views. Then we train the network in a contrastive semi-supervised manner by reconstructing label embeddings for both reliable supervision of clean data and sufficient exploration of noisy data. On the other hand, inspired by recent pruning techniques, DIOR divides the parameters in the hashing network into crucial parameters and non-crucial parameters, and then optimizes them separately to reduce the overfitting of noisy data. Extensive experiments on four popular benchmark datasets demonstrate the effectiveness of DIOR. Haixin Wang 0003, Huiyu Jiang, Jinan Sun, Shikun Zhang, Chong Chen 0002, Xian-Sheng Hua 0001, Xiao Luo 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2024 | Look Into Gradients: Learning Compact Hash Codes for Out-of-Distribution RetrievalabstractHashing aims to compress raw data into compact binary descriptors, which has drawn increasing interest for efficient large-scale image retrieval. Current deep hashing often employs evaluation protocols where usually query data and training data are from similar distributions. However, more realistic evaluations should take into account a broad spectrum of distribution shifts with varying degrees. Therefore, we study the problem of out-of-distribution generalization in image retrieval, which seeks to learn a retrieval model from a source domain and generalize to unseen target domains. However, this problem is challenging owing to data scarcity in target domains and the potential overfitting of domain-specific patterns. Here, we propose a novel hashing model namedLooking-into-gradients (LOG) for image retrieval under out-of-distribution shifts, which comprehensively explores gradients for both data generation and model optimization. Specifically, to overcome data deficiency in target domains, we formalize the worst-case problem to generate challenging virtue samples via adversarial gradient ascend. Besides, to further enhance model generalization capability, we not only identify non-crucial parameters with minor gradients and values and shrink them to zero, but also modify the inconsistent gradients across domains to prevent learning domain-specific patterns. Extensive experiments on various datasets demonstrate that LOG outperforms state-of-the-art methods by up to 8.54%. Haixin Wang 0003, Xinlong Yang, Jinan Sun, Shikun Zhang, Chong Chen 0002, Xian-Sheng Hua 0001, Xiao Luo 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2024 | Leveraging Imitation Learning on Pose Regulation Problem of a Robotic FishabstractIn this article, the pose regulation control problem of a robotic fish is investigated by formulating it as a Markov decision process (MDP). Such a typical task that requires the robot to arrive at the desired position with the desired orientation remains a challenge, since two objectives (position and orientation) may be conflicted during optimization. To handle the challenge, we adopt the sparse reward scheme, i.e., the robot will be rewarded if and only if it completes the pose regulation task. Although deep reinforcement learning (DRL) can achieve such an MDP with sparse rewards, the absence of immediate reward hinders the robot from efficient learning. To this end, we propose a novel imitation learning (IL) method that learns DRL-based policies from demonstrations with inverse reward shaping to overcome the challenge raised by extremely sparse rewards. Moreover, we design a demonstrator to generate various trajectory demonstrations based on one simple example from a nonexpert helper, which greatly reduces the time consumption of collecting robot samples. The simulation results evaluate the effectiveness of our proposed demonstrator and the state-of-the-art (SOTA) performance of our proposed IL method. Furthermore, we deploy the trained IL policy on a physical robotic fish to perform pose regulation in a swimming tank without/with external disturbances. The experimental results verify the effectiveness and robustness of our proposed methods in real world. Therefore, we believe this article is a step forward in the field of biomimetic underwater robot learning. Lu Yue, Chen Wang 0005, Jinan Sun, Shikun Zhang, Airong Wei, Guangming Xie |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Reviewing Labels: Label Graph Network with Top-k Prediction Set for Relation ExtractionabstractThe typical way for relation extraction is fine-tuning large pre-trained language models on task-specific datasets, then selecting the label with the highest probability of the output distribution as the final prediction. However, the usage of the Top-k prediction set for a given sample is commonly overlooked. In this paper, we first reveal that the Top-k prediction set of a given sample contains useful information for predicting the correct label. To effectively utilizes the Top-k prediction set, we propose Label Graph Network with Top-k Prediction Set, termed as KLG. Specifically, for a given sample, we build a label graph to review candidate labels in the Top-k prediction set and learn the connections between them. We also design a dynamic k selection mechanism to learn more powerful and discriminative relation representation. Our experiments show that KLG achieves the best performances on three relation extraction datasets. Moreover, we observe thatKLG is more effective in dealing with long-tailed classes. Bo Li 0099, Wei Ye 0004, Shikun Zhang |
AAAI | 4 |
| 2023 | Sequence Generation with Label Augmentation for Relation ExtractionabstractSequence generation demonstrates promising performance in recent information extraction efforts, by incorporating large-scale pre-trained Seq2Seq models. This paper investigates the merits of employing sequence generation in relation extraction, finding that with relation names or synonyms as generation targets, their textual semantics and the correlation (in terms of word sequence pattern) among them affect model performance. We then propose Relation Extraction with Label Augmentation (RELA), a Seq2Seq model with automatic label augmentation for RE. By saying label augmentation, we mean prod semantically synonyms for each relation name as the generation target. Besides, we present an in-depth analysis of the Seq2Seq model's behavior when dealing with RE. Experimental results show that RELA achieves competitive results compared with previous methods on four RE datasets. Bo Li 0099, Dingyao Yu, Wei Ye 0004, Shikun Zhang |
AAAI | 5 |
| 2023 | Vision Language Pre-training by Contrastive Learning with Cross-Modal Similarity RegulationabstractCross-modal contrastive learning in vision language pretraining (VLP) faces the challenge of (partial) false negatives.In this paper, we study this problem from the perspective of Mutual Information (MI) optimization.It is common sense that InfoNCE loss used in contrastive learning will maximize the lower bound of MI between anchors and their positives, while we theoretically prove that MI involving negatives also matters when noises commonly exist.Guided by a more general lower bound form for optimization, we propose a contrastive learning strategy regulated by progressively refined cross-modal similarity, to more accurately optimize MI between an image/text anchor and its negative texts/images instead of improperly minimizing it.Our method performs competitively on four downstream cross-modal tasks and systematically balances the beneficial and harmful effects of (partial) false negative samples under theoretical guidance. Chaoya Jiang, Wei Ye 0004, Haiyang Xu 0001, Songfang Huang, Fei Huang 0002, Shikun Zhang |
ACL (1) | 6 |
| 2023 | 3D Registration with Maximal CliquesabstractAs a fundamental problem in computer vision, 3D point cloud registration (PCR) aims to seek the optimal pose to align a point cloud pair. In this paper, we present a 3D registration method with maximal cliques (MAC). The key insight is to loosen the previous maximum clique constraint, and mine more local consensus information in a graph for accurate pose hypotheses generation: 1) A compatibility graph is constructed to render the affinity relationship between initial correspondences. 2) We search for maximal cliques in the graph, each of which represents a consensus set. We perform node-guided clique selection then, where each node corresponds to the maximal clique with the greatest graph weight. 3) Transformation hypotheses are computed for the selected cliques by the SVD algorithm and the best hypothesis is used to perform registration. Extensive experiments on U3M, 3DMatch, 3DLoMatch and KITTI demonstrate that MAC effectively increases registration accuracy, outperforms various state-of-the-art methods and boosts the performance of deep-learned methods. MAC combined with deep-learned methods achieves state-of-the-art registration recall of 95.7% /78.9% on 3DMatch /3DLoMatch. Xiyu Zhang 0001, Jiaqi Yang 0002, Shikun Zhang, Yanning Zhang 0001 |
CVPR | 3 |
| 2023 | SDRNet: Shape Decoupled Regression Network for 3d face ReconstructionabstractIn the field of computer vision, 3D face reconstruction from single-view images is a long-standing and challenging problem. Following the popular 3DMM-based reconstruction framework, recent works show great concerns about exploring discriminative information of identity and expression for shape regression commonly in coupling ways. Actually, identity and expression information may contribute differently in explaining the intrinsic shape of faces, and the former is inferior to the latter in explaining the great facial shape variations caused by extreme expression. In this paper, we propose a Shape Decoupled Regression Network (SDRNet) consisting of identity-focused branch and expression-focused branch with focused criteria for representation learning, which interact with the union branch to achieve the final 3DMM parameters regression for improved shape reconstruction. In SDRNet, the focused criteria estimate the 3D vertex prediction loss, while the predicted 3D shape is reconstructed only using the predicted parameter of identity or expression and introducing the ground-truth parameters of the left two. Extensive experiments on the challenging AFLW2000-3D and AFLW datasets demonstrate advanced performance in 3D face reconstruction and face alignment. Shikun Zhang, Fengyi Song, Ming Yang 0014 |
ICASSP | 1 |
| 2023 | BUS : Efficient and Effective Vision-language Pre-training with Bottom-Up Patch SummarizationabstractVision Transformer (ViT) based Vision-Language Pre-training (VLP) models have demonstrated impressive performance in various tasks. However, the lengthy visual token sequences fed into ViT can lead to training inefficiency and ineffectiveness. Existing efforts address the challenge by either bottom-level patch extraction in the ViT backbone or top-level patch abstraction outside, not balancing training efficiency and effectiveness well. Inspired by text summarization in natural language processing, we propose a Bottom-Up Patch Summarization approach named BUS, coordinating bottom-level extraction and top-level abstraction to learn a concise summary of lengthy visual token sequences efficiently. Specifically, We incorporate a Text-Semantics-Aware Patch Selector (TSPS) into the ViT backbone to perform a coarse-grained visual token extraction and then attach a flexible Transformer-based Patch Abstraction Decoder (PAD) upon the backbone for top-level visual abstraction. This bottom-up collaboration enables our BUS to yield high training efficiency while maintaining or even improving effectiveness. We evaluate our approach on various visual-language understanding and generation tasks and show competitive downstream task performance while boosting the training efficiency by 50%. Additionally, our model achieves state-of-the-art performance on many downstream tasks by increasing input image resolution without increasing computational costs over baselines. Chaoya Jiang, Haiyang Xu 0001, Wei Ye 0004, Qinghao Ye, Chenliang Li 0003, Ming Yan 0008, Bin Bi, Shikun Zhang, Fei Huang 0002, Songfang Huang |
ICCV | 8 |
| 2023 | Hierarchical Prior Mining for Non-local Multi-View StereoabstractAs a fundamental problem in computer vision, multi-view stereo (MVS) aims at recovering the 3D geometry of a target from a set of 2D images. Recent advances in MVS have shown that it is important to perceive non-local structured information for recovering geometry in low-textured areas. In this work, we propose a Hierarchical Prior Mining for Non-local Multi-View Stereo (HPM-MVS). The key characteristics are the following techniques that exploit non-local information to assist MVS: 1) A Non-local Extensible Sampling Pattern (NESP), which is able to adaptively change the size of sampled areas without becoming snared in locally optimal solutions. 2) A new approach to leverage non-local reliable points and construct a planar prior model based on K-Nearest Neighbor (KNN), to obtain potential hypotheses for the regions where prior construction is challenging. 3) A Hierarchical Prior Mining (HPM) framework, which is used to mine extensive non-local prior information at different scales to assist 3D model recovery, this strategy can achieve a considerable balance between the reconstruction of details and low-textured areas. Experimental results on the ETH3D and Tanks & Temples have verified the superior performance and strong generalization capability of our method. Our code will be available at https://github.com/CLinvx/HPM-MVS. Chunlin Ren, Qingshan Xu 0001, Shikun Zhang, Jiaqi Yang 0002 |
ICCV | 3 |
| 2023 | Prototypical Mixing and Retrieval-based Refinement for Label Noise-resistant Image RetrievalabstractLabel noise is pervasive in real-world applications, which influences the optimization of neural network models. This paper investigates a realistic but understudied problem of image retrieval under label noise, which could lead to severe overfitting or memorization of noisy samples during optimization. Moreover, identifying noisy samples correctly is still a challenging problem for retrieval models. In this paper, we propose a novel approach called Prototypical Mixing and Retrieval-based Refinement (TITAN) for label noise-resistant image retrieval, which corrects label noise and mitigates the effects of the memorization simultaneously. Specifically, we first characterize numerous prototypes with Gaussian distributions in the hidden space, which would direct the Mixing procedure in providing synthesized samples. These samples are fed into a similarity learning framework with varying emphasis based on the prototypical structure to learn semantics with reduced overfitting. In addition, we retrieve comparable samples for each prototype from simple to complex, which refine noisy samples in an accurate and class-balanced manner. Comprehensive experiments on five benchmark datasets demonstrate the superiority of our proposed TITAN compared with various competing baselines. Xinlong Yang, Haixin Wang 0003, Jinan Sun, Shikun Zhang, Chong Chen 0002, Xian-Sheng Hua 0001, Xiao Luo 0001 |
ICCV | 4 |
| 2023 | Assessing and Improving Dataset and Evaluation Methodology in Deep Learning for Code Clone DetectionabstractCode clone detection is a task that identifies whether two code snippets are semantically identical. In recent years, deep learning models have shown high performance in detecting Type-3 and Type-4 code clones, and received increasing attention from the research community. However, compared with the attention given to the model design by the researchers, there is little research work on the quality of the datasets and the evaluation methodology (the way of dividing the dataset into training set and test set), which poses a challenge to the credibility of deep learning models.In this paper, we conduct experiments to evaluate the performance of the existing state-of-the-art models in multi-perspectives. At the same time, we release two new datasets for code clone detection, namely ConBigCloneBench and Google-CodeJam2 based on the existing datasets BigCloneBench and GoogleCodeJam, respectively. Our experiments show that the performance of the same model decreases up to 0.5 F1 score (from 0.9 to 0.4) on different evaluation perspectives and datasets, and the performance of some models is only similar to the simple MLP model. We analyze reasons for the performance decline further, and provide suggestions for future research to improve the performance of deep learning models from multi-perspectives. Shikun Zhang |
ISSRE | 3 |
| 2023 | COPA : Efficient Vision-Language Pre-training through Collaborative Object- and Patch-Text AlignmentabstractVision-Language Pre-training (VLP) methods based on object detection enjoy the rich knowledge of fine-grained object-text alignment but at the cost of computationally expensive inference. Recent Visual-Transformer (ViT)-based approaches circumvent this issue while struggling with long visual sequences without detailed cross-modal alignment information. This paper introduces a ViT-based VLP technique that efficiently incorporates object information through a novel patch-text alignment mechanism. Specifically, we convert object-level signals into patch-level ones and devise a Patch-Text Alignment pre-training task (PTA) to learn a text-aware patch detector. By using off-the-shelf delicate object annotations in 5% training images, we jointly train PTA with other conventional VLP objectives in an end-to-end manner, bypassing the high computational cost of object detection and yielding an effective patch detector that accurately detects text-relevant patches, thus considerably reducing patch sequences and accelerating computation within the ViT backbone. Our experiments on a variety of widely-used benchmarks reveal that our method achieves a speedup of nearly 88% compared to prior VLP models while maintaining competitive or superior performance on downstream tasks with similar model size and data scale. Chaoya Jiang, Haiyang Xu 0001, Wei Ye 0004, Qinghao Ye, Chenliang Li 0003, Ming Yan 0008, Bin Bi, Shikun Zhang, Fei Huang 0002, Ji Zhang 0011 |
ACM Multimedia | 8 |
| 2023 | Parameter-efficient Tuning of Large-scale Multimodal Foundation ModelabstractDriven by the progress of large-scale pre-training, parameter-efficient transfer learning has gained immense popularity across different subfields of Artificial Intelligence. The core is to adapt the model to downstream tasks with only a small set of parameters. Recently, researchers have leveraged such proven techniques in multimodal tasks and achieve promising results. However, two critical issues remain unresolved: how to further reduce the complexity with lightweight design and how to boost alignment between modalities under extremely low parameters. In this paper, we propose A gracefUl pRompt framewOrk for cRoss-modal trAnsfer (AURORA) to overcome these challenges. Considering the redundancy in existing architectures, we first utilize the mode approximation to generate 0.1M trainable parameters to implement the multimodal parameter-efficient tuning, which explores the low intrinsic dimension with only 0.04% parameters of the pre-trained model. Then, for better modality alignment, we propose the Informative Context Enhancement and Gated Query Transformation module under extremely few parameters scenes. A thorough evaluation on six cross-modal benchmarks shows that it not only outperforms the state-of-the-art but even outperforms the full fine-tuning approach. Our code is available at: https://github.com/WillDreamer/Aurora. Haixin Wang 0003, Xinlong Yang, Jianlong Chang, Dian Jin 0004, Jinan Sun, Shikun Zhang, Xiao Luo 0001, Qi Tian 0001 |
NeurIPS | 6 |
| 2023 | IDEA: An Invariant Perspective for Efficient Domain Adaptive Image RetrievalabstractIn this paper, we investigate the problem of unsupervised domain adaptive hashing, which leverage knowledge from a label-rich source domain to expedite learning to hash on a label-scarce target domain. Although numerous existing approaches attempt to incorporate transfer learning techniques into deep hashing frameworks, they often neglect the essential invariance for adequate alignment between these two domains. Worse yet, these methods fail to distinguish between causal and non-causal effects embedded in images, rendering cross-domain retrieval ineffective. To address these challenges, we propose an Invariance-acquired Domain AdaptivE HAshing (IDEA) model. Our IDEA first decomposes each image into a causal feature representing label information, and a non-causal feature indicating domain information. Subsequently, we generate discriminative hash codes using causal features with consistency learning on both source and target domains. More importantly, we employ a generative model for synthetic samples to simulate the intervention of various non-causal effects, ultimately minimizing their impact on hash codes for domain invariance. Comprehensive experiments conducted on benchmark datasets validate the superior performance of our IDEA compared to a variety of competitive baselines. Haixin Wang 0003, Hao Wu 0094, Jinan Sun, Shikun Zhang, Chong Chen 0002, Xian-Sheng Hua 0001, Xiao Luo 0001 |
NeurIPS | 4 |
| 2023 | DANCE: Learning A Domain Adaptive Framework for Deep HashingabstractThis paper studies unsupervised domain adaptive hashing, which aims to transfer a hashing model from a label-rich source domain to a label-scarce target domain. Current state-of-the-art approaches generally resolve the problem by integrating pseudo-labeling and domain adaptation techniques into deep hashing paradigms. Nevertheless, they usually suffer from serious class imbalance in pseudo-labels and suboptimal domain alignment caused by the neglection of the intrinsic structures of two domains. To address this issue, we propose a novel method named unbiaseD duAl hashiNg Contrastive lEarning (DANCE) for domain adaptive image retrieval. The core of our DANCE is to perform contrastive learning on hash codes from both instance level and prototype level. To begin, DANCE utilizes label information to guide instance-level hashing contrastive learning in the source domain. To generate unbiased and reliable pseudo-labels for semantic learning in the target domain, we uniformly select samples around each label embedding in the Hamming space. A momentum-update scheme is also utilized to smooth the optimization process. Additionally, we measure the semantic prototype representations in both source and target domains and incorporate them into a domain-aware prototype-level contrastive learning paradigm, which enhances domain alignment in the Hamming space while maximizing the model capacity. Experimental results on a number of well-known domain adaptive retrieval benchmarks validate the effectiveness of our proposed DANCE compared to a variety of competing baselines in different settings. Haixin Wang 0003, Jinan Sun, Shikun Zhang, Chong Chen 0002, Xian-Sheng Hua 0001, Xiao Luo 0001 |
WWW | 4 |
| 2023 | Toward Effective Domain Adaptive RetrievalabstractThis paper studies the problem of unsupervised domain adaptive hashing, which is less-explored but emerging for efficient image retrieval, particularly for cross-domain retrieval. This problem is typically tackled by learning hashing networks with pseudo-labeling and domain alignment techniques. Nevertheless, these approaches usually suffer from overconfident and biased pseudo-labels and inefficient domain alignment without sufficiently exploring semantics, thus failing to achieve satisfactory retrieval performance. To tackle this issue, we present PEACE, a principled framework which holistically explores semantic information in both source and target data and extensively incorporates it for effective domain alignment. For comprehensive semantic learning, PEACE leverages label embeddings to guide the optimization of hash codes for source data. More importantly, to mitigate the effects of noisy pseudo-labels, we propose a novel method to holistically measure the uncertainty of pseudo-labels for unlabeled target data and progressively minimize them through alternative optimization under the guidance of the domain discrepancy. Additionally, PEACE effectively removes domain discrepancy in the Hamming space from two views. In particular, it not only introduces composite adversarial learning to implicitly explore semantic information embedded in hash codes, but also aligns cluster semantic centroids across domains to explicitly exploit label information. Experimental results on several popular domain adaptive retrieval benchmarks demonstrate the superiority of our proposed PEACE compared with various state-of-the-art methods on both single-domain and cross-domain retrieval tasks. Our source codes are available at https://github.com/WillDreamer/PEACE. Haixin Wang 0003, Jinan Sun, Xiao Luo 0001, Shikun Zhang, Chong Chen 0002, Xian-Sheng Hua 0001 |
IEEE Trans. Image Process. | 5 |
| 2022 | Frequency-Aware Contrastive Learning for Neural Machine TranslationabstractLow-frequency word prediction remains a challenge in modern neural machine translation (NMT) systems. Recent adaptive training methods promote the output of infrequent words by emphasizing their weights in the overall training objectives. Despite the improved recall of low-frequency words, their prediction precision is unexpectedly hindered by the adaptive objectives. Inspired by the observation that low-frequency words form a more compact embedding space, we tackle this challenge from a representation learning perspective. Specifically, we propose a frequency-aware token-level contrastive learning method, in which the hidden state of each decoding step is pushed away from the counterparts of other target words, in a soft contrastive way based on the corresponding word frequencies. We conduct experiments on widely used NIST Chinese-English and WMT14 English-German translation tasks. Empirical results show that our proposed methods can not only significantly improve the translation quality but also enhance lexical diversity and optimize word representation space. Further investigation reveals that, comparing with related adaptive training strategies, the superiority of our method on low-frequency word prediction lies in the robustness of token-level recall across different frequencies without sacrificing precision. Tong Zhang 0001, Wei Ye 0004, Baosong Yang, Long Zhang 0012, Xingzhang Ren, Dayiheng Liu, Jinan Sun, Shikun Zhang, Haibo Zhang 0013 |
AAAI | 8 |
| 2022 | Label Smoothing for Text MiningabstractCurrent text mining models are trained with 0-1 hard label that indicates whether an instance belongs to a class, ignoring rich information of the relevance degree. Soft label, which involved each label of varying degrees than the hard label, is considered more suitable for describing instances. The process of generating soft labels from hard labels is defined as label smoothing (LS). Classical LS methods focus on universal data mining tasks so that they ignore the valuable text features in text mining tasks. This paper presents a novel keyword-based LS method to automatically generate soft labels from hard labels via exploiting the relevance between labels and text instances. Generated soft labels are then incorporated into existing models as auxiliary targets during the training stage, capable of improving models without adding any extra parameters. Results of extensive experiments on text classification and large-scale text retrieval datasets demonstrate that soft labels generated by our method contain rich knowledge of text features, improving the performance of corresponding models under both balanced and unbalanced settings. Peiyang Liu, Xiangyu Xi, Wei Ye 0004, Shikun Zhang |
COLING | 4 |
| 2022 | Exploiting Hybrid Semantics of Relation Paths for Multi-hop Question Answering over Knowledge GraphsabstractAnswering natural language questions on knowledge graphs (KGQA) remains a great challenge in terms of understanding complex questions via multi-hop reasoning. Previous efforts usually exploit large-scale entity-related text corpus or knowledge graph (KG) embeddings as auxiliary information to facilitate answer selection. However, the rich semantics implied in off-the-shelf relation paths between entities is far from well explored. This paper proposes improving multi-hop KGQA by exploiting relation paths’ hybrid semantics. Specifically, we integrate explicit textual information and implicit KG structural features of relation paths based on a novel rotate-and-scale entity link prediction framework. Extensive experiments on three existing KGQA datasets demonstrate the superiority of our method, especially in multi-hop scenarios. Further investigation confirms our method’s systematical coordination between questions and relation paths to identify answer entities. Zile Qiao, Wei Ye 0004, Tong Zhang 0001, Tong Mo, Shikun Zhang |
COLING | 6 |
| 2022 | TRIPS: Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch SelectionabstractVision Transformers (ViTs) have been widely used in large-scale Vision and Language Pretraining (VLP) models.Though previous VLP works have proved the effectiveness of ViTs, they still suffer from computational efficiency brought by the long visual sequence.To tackle this problem, in this paper, we propose an efficient vision-and-language pre-training model with Text-Relevant Image Patch Selection, namely TRIPS, which reduces the visual sequence progressively with a text-guided patchselection layer in the visual backbone for efficient training and inference.The patchselection layer can dynamically compute textdependent visual attention to identify the attentive image tokens with text guidance and fuse inattentive ones in an end-to-end manner.Meanwhile, TRIPS does not introduce extra parameters to ViTs.Experimental results on a variety of popular benchmark datasets demonstrate that TRIPS gain a speedup of 40% over previous similar VLP models, yet with competitive or better downstream task performance. Chaoya Jiang, Haiyang Xu 0001, Chenliang Li 0003, Ming Yan 0008, Wei Ye 0004, Shikun Zhang, Bin Bi, Songfang Huang |
EMNLP | 6 |
| 2022 | A Lightweight Network with Multi-Stage Feature Fusion Module for Single-View 3d Face Reconstructionabstract3D face reconstruction has attracted great attentions of researchers from both academic and industry for its potential application in many scenarios such as face alignment and recognition across large poses. 3D Morphable Model which reconstructs a 3D face through basis coefficients prediction, is usually adopted as the typical parametric framework for 3D face and is suitable to combine with deep learning. Existing cascade regression method predicts coefficients by multiple iterations, which is time-consuming. In this paper, we propose an efficient and end-to-end method for single-view 3D face reconstruction. We build a lightweight network based on mobile blocks with faster speed for parameter extraction and smaller model size. Especially, a multi-stage feature fusion module is designed for enhancing the end-to-end learning. To match the setting of input image size, we updated the pose label of images under various sizes in training dataset before training. Extensive experiments on challenging datasets validate the efficiency of our method for both 3D face reconstruction and face alignment. Shikun Zhang, Fengyi Song, Ming Yang 0014 |
ICIP | 2 |
| 2022 | Exploring Occlusion-Sensitive Deep Network for Single-View 3D Face ReconstructionabstractRecovering 3D geometry from a single-view 2D face image is an ill-posed task full of various challenges, especially under occlusion conditions commonly seen with large poses, while partial facial information missing makes the burden much heavier. Although existing methods could solve this problem in the end-to-end fashion, they still have limited performance in the occluded scenes. It is intuitively for many methods to depress the influence of those occluded regions for reconstruction. But we propose an occlusion-sensitive weighting mechanism for balancing the contributions among occluded and non-occluded regions. Meanwhile, considering no dataset contains various occlusions for learning, the data augmentation technique is exploited to expand the training dataset, which further facilitates the learning of the occlusion-sensitive deep network. Extensive experiments on two challenging datasets validate the advanced performance of our method for both 3D face reconstruction and face alignment.1 Shikun Zhang, Fengyi Song, Ming Yang 0014 |
ICIP | 2 |
| 2022 | 3Rs: Data Augmentation Techniques Using Document Contexts For Low-Resource Chinese Named Entity RecognitionabstractWith recent advances of neural networks and pre-training techniques, Chinese Named Entity Recognition (NER) has achieved great progress in recent years. However, NER systems still have the problem of generalization ability issues due to lack of annotated data, and current NER models mostly consider input sentences individually, which prevent models from further exploiting cross-sentence document context in training. With regard of these problems, this paper present new insights into Chinese NER and propose 3Rs: three data augmentation methods incorporating document-level information for NER through random concatenating, random swapping and random erasing, which are inspired by some multi-sample data augmentation techniques in computer vision fields, aiming to reorganize the composition of training sentences, and generate more training examples with less human efforts. We conduct extensive experiments on two Chinese datasets, and introduce a two-level attacking method to audit robustness performance. Our experiment results show that even the best model can obtain a better accuracy and robustness, especially for smaller training sets, therefore alleviating performance bottlenecks on low-resource conditions. Zheyu Ying, Rui Xie 0003, Guochang Wen, Xueyang Liu, Shikun Zhang |
IJCNN | 7 |
| 2022 | Low-Resources Project-Specific Code SummarizationabstractCode summarization generates brief natural language descriptions of source code pieces, which can assist developers in understanding code and reduce documentation workload. Recent neural models on code summarization are trained and evaluated on large-scale multi-project datasets consisting of independent code-summary pairs. Despite the technical advances, their effectiveness on a specific project is rarely explored. In practical scenarios, however, developers are more concerned with generating high-quality summaries for their working projects. And these projects may not maintain sufficient documentation, hence having few historical code-summary pairs. To this end, we investigate low-resource project-specific code summarization, a novel task more consistent with the developers’ requirements. To better characterize project-specific knowledge with limited training samples, we propose a meta transfer learning method by incorporating a lightweight fine-tuning mechanism into a meta-learning framework. Experimental results on nine real-world projects verify the superiority of our method over alternative ones and reveal how the project-specific knowledge is learned. Rui Xie 0003, Tianxiang Hu, Wei Ye 0004, Shikun Zhang |
ASE | 4 |
| 2022 | HEART: Towards Effective Hash Codes under Label NoiseabstractHashing, which encodes raw data into compact binary codes, has grown in popularity for large-scale image retrieval due to its storage and computation efficiency. Although deep supervised hashing has lately shown promising performance, they mostly assume that the semantic labels of training data are ideally noise-free, which is often unrealistic in real-world applications. In this paper, considering the practical application, we focus on the problem of learning to hash with label noise and propose a novel method called HEART to address the problem. HEART is a holistic framework which explores latent semantic distributions to select both clean samples and pairs of high confidence for mitigating the impacts of label noise. From a statistical perspective, our HEART characterizes each image by its multiple augmented views that can be considered as examples from its latent distribution and then calculates semantic distances between images using energy distances between their latent distributions. With semantic distances, we can select confident similar pairs to guide hashing contrastive learning for high-quality hash codes. Moreover, to prevent the memorization of noisy examples, we propose a novel strategy to identify clean samples which have small variations of losses on the latent distributions and train the network on clean samples using a pointwise loss. Experimental results on several popular benchmark datasets demonstrate the effectiveness of our HEART compared with a wide range of baselines. Jinan Sun, Haixin Wang 0003, Xiao Luo 0001, Shikun Zhang, Chong Chen 0002, Xian-Sheng Hua 0001 |
ACM Multimedia | 4 |
| 2022 | Museformer: Transformer with Fine- and Coarse-Grained Attention for Music GenerationabstractSymbolic music generation aims to generate music scores automatically. A recent trend is to use Transformer or its variants in music generation, which is, however, suboptimal, because the full attention cannot efficiently model the typically long music sequences (e.g., over 10,000 tokens), and the existing models have shortcomings in generating musical repetition structures. In this paper, we propose Museformer, a Transformer with a novel fine- and coarse-grained attention for music generation. Specifically, with the fine-grained attention, a token of a specific bar directly attends to all the tokens of the bars that are most relevant to music structures (e.g., the previous 1st, 2nd, 4th and 8th bars, selected via similarity statistics); with the coarse-grained attention, a token only attends to the summarization of the other bars rather than each token of them so as to reduce the computational cost. The advantages are two-fold. First, it can capture both music structure-related correlations via the fine-grained attention, and other contextual information via the coarse-grained attention. Second, it is efficient and can model over 3X longer music sequences compared to its full-attention counterpart. Both objective and subjective experimental results demonstrate its ability to generate long music sequences with high quality and better structures. Botao Yu, Peiling Lu, Rui Wang 0028, Xu Tan 0003, Wei Ye 0004, Shikun Zhang, Tao Qin 0001, Tie-Yan Liu |
NeurIPS | 7 |
| 2022 | How Usable Are iOS App Privacy Labels?abstractStandardized privacy labels that succinctly summarize those data practices that people are most commonly concerned about offer the promise of providing users with more effective privacy notices than full-length privacy policies. With their introduction by Apple in iOS 14 and Google’s recent adoption in its Play Store, mobile app privacy labels are for the first time available at scale to users. We report the first indepth interview study with 24 lay iPhone users to investigate their experiences, understanding, and perceptions of Apple’s privacy labels. We uncovered misunderstandings of and dissatisfaction with the iOS privacy labels that hinder their effectiveness, including confusing structure, unfamiliar terms, and disconnection from permission settings and controls. We identify areas where app privacy labels might be improved and propose suggestions to address shortcomings to make them more understandable, usable, and useful. Shikun Zhang, Yuanyuan Feng, Yaxing Yao, Lorrie Faith Cranor, Norman M. Sadeh |
Proc. Priv. Enhancing Technol. | 1 |
| 2022 | TencentCLS: The Cloud Log Service with High Query PerformancesabstractWith the trend of cloud computing, the cloud log service is becoming increasingly important, as it plays a critical role in tasks such as root cause analysis, service monitoring and security audition. To meet these needs, we provide Tencent Cloud Log Service (TencentCLS), a one-stop solution for log collection, storage, analysis and dumping. It currently hosts more than a million tenants, of which the largest ones can generate up to PB-level logs per day. The most important challenge that TencentCLS faces is to support both low-latency and resource-efficient queries on such large quantities of log data. To address that challenge, we propose a novel search engine based upon Lucene. The system features a novel procedure for querying logs within a time range, an indexing technique for the time field, as well as optimized query algorithms dedicated to multiple critical and common query types. As a result, the search engine at TencentCLS gains significant performance improvements against Lucene. It achieves 20x performance increase with standard queries, and 10x performance increase with histogram queries in massive log query scenarios. In addition, TencentCLS also supports storing and querying with microsecond-level time precision, as well as the microsecond-level time order preservation capability. Muzhi Yu, Zhaoxiang Lin, Jinan Sun, Runyun Zhou, Guoqiang Jiang, Shikun Zhang |
Proc. VLDB Endow. | 7 |
| 2022 | From Simulation to Reality: A Learning Framework for Fish-Like Robots to Perform Control TasksabstractThe fish-like robot is one of the typical underwater robots, which has the advantage of high maneuverability with low noise due to its bioinspired structure and biomimetic locomotion. However, it is challenging to efficiently design motion controllers for such robots to achieve satisfactory performance on specific control tasks in the real underwater environment, since the complex fluid-structure interaction exists during their swimming and exact dynamic models are absent. In this article, we propose a learning framework, incorporating a simulation system and a training methodology, to autonomously and fast train in simulation to create control policies that are capable of directly applying to a type of physical fish-like robots to perform motion control tasks. First, we construct a simulation system combining a data-driven environment and a computational fluid dynamics (CFD)-based environment, thus well balancing the simulation accuracy and the calculation speed. Second, we design a training methodology to train deep reinforcement learning (DRL)-based policies for the robot in our constructed simulation system to perform a specific control task. Then, we use two typical motion control tasks to verify our proposed framework. One is the path-following control task, which is a one-objective problem with dense rewards, while the other is the pose control task which is a two-objective problem with sparse rewards. For each task, the DRL-based control policy trained by our learning framework is directly deployed on the physical fish-like robot to perform the task in the real world. Experimental results show that the policies trained in simulation still work well in the real world, and perform even better in terms of control accuracy and stability compared with the traditional control methods, thus demonstrating the effectiveness of our learning framework. Runyu Tian, Hongqi Yang, Chen Wang 0005, Jinan Sun, Shikun Zhang, Guangming Xie |
IEEE Trans. Robotics | 6 |
| 2021 | Multi-view Inference for Relation Extraction with Uncertain KnowledgeabstractKnowledge graphs (KGs) are widely used to facilitate relation extraction (RE) tasks. While most previous RE methods focus on leveraging deterministic KGs, uncertain KGs, which assign a confidence score for each relation instance, can provide prior probability distributions of relational facts as valuable external knowledge for RE models. This paper proposes to exploit uncertain knowledge to improve relation extraction. Specifically, we introduce ProBase, an uncertain KG that indicates to what extent a target entity belongs to a concept, into our RE architecture. We then design a novel multi-view inference framework to systematically integrate local context and global knowledge across three views: mention-, entity- and concept-view. The experiment results show that our model achieves competitive performances on both sentence- and document-level relation extraction, which verifies the effectiveness of introducing uncertain knowledge and the multi-view inference framework that we design. Bo Li 0099, Wei Ye 0004, Canming Huang, Shikun Zhang |
AAAI | 4 |
| 2021 | SongMASS: Automatic Song Writing with Pre-training and Alignment ConstraintabstractAutomatic song writing aims to compose a song (lyric and/or melody) by machine, which is an interesting topic in both academia and industry. In automatic song writing, lyric-to-melody generation and melody-to-lyric generation are two important tasks, both of which usually suffer from the following challenges: 1) the paired lyric and melody data are limited, which affects the generation quality of the two tasks, considering a lot of paired training data are needed due to the weak correlation between lyric and melody; 2) Strict alignments are required between lyric and melody, which relies on specific alignment modeling. In this paper, we propose SongMASS to address the above challenges, which leverages masked sequence to sequence (MASS) pre-training and attention based alignment modeling for lyric-to-melody and melody-to-lyric generation. Specifically, 1) we extend the original sentence-level MASS pre-training to song level to better capture long contextual information in music, and use a separate encoder and decoder for each modality (lyric or melody); 2) we leverage sentence-level attention mask and token-level attention constraint during training to enhance the alignment between lyric and melody. During inference, we use a dynamic programming strategy to obtain the alignment between each word/syllable in lyric and note in melody. We pre-train SongMASS on unpaired lyric and melody datasets, and both objective and subjective evaluations demonstrate that SongMASS generates lyric and melody with significantly better quality than the baseline method. Zhonghao Sheng, Kaitao Song, Xu Tan 0003, Yi Ren 0006, Wei Ye 0004, Shikun Zhang, Tao Qin 0001 |
AAAI | 6 |
| 2021 | Capturing Event Argument Interaction via A Bi-Directional Entity-Level Recurrent DecoderabstractXi Xiangyu, Wei Ye, Shikun Zhang, Quanxiu Wang, Huixing Jiang, Wei Wu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Xiangyu Xi, Wei Ye 0004, Shikun Zhang, Quanxiu Wang, Huixing Jiang, Wei Wu 0014 |
ACL/IJCNLP (1) | 3 |
| 2021 | Unsupervised Out-of-Domain Detection via Pre-trained TransformersabstractKeyang Xu, Tongzheng Ren, Shikun Zhang, Yihao Feng, Caiming Xiong. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Keyang Xu, Tongzheng Ren, Shikun Zhang, Yihao Feng, Caiming Xiong |
ACL/IJCNLP (1) | 3 |
| 2021 | Point, Disambiguate and Copy: Incorporating Bilingual Dictionaries for Neural Machine TranslationabstractTong Zhang, Long Zhang, Wei Ye, Bo Li, Jinan Sun, Xiaoyu Zhu, Wen Zhao, Shikun Zhang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Tong Zhang 0001, Long Zhang 0012, Wei Ye 0004, Bo Li 0099, Jinan Sun, Shikun Zhang |
ACL/IJCNLP (1) | 8 |
| 2021 | Distilling Knowledge from BERT into Simple Fully Connected Neural Networks for Efficient Vertical RetrievalabstractDistilled BERT models are more suitable for efficient vertical retrieval in online sponsored vertical search with low-latency requirements than BERT due to fewer parameters and faster inference. Unfortunately, most of these models are still far from ideal inference speed. This paper presents a novel and effective method to distill knowledge from BERT into simple fully connected neural networks (FNN). Results of extensive experiments on English and Chinese datasets demonstrate that our method achieves comparable results with existing distilled BERT models while the inference is accelerated by more than ten times. We have successfully applied our method on our online sponsored vertical search engine and get remarkable improvements. Peiyang Liu, Lin Wang 0106, Wei Ye 0004, Xiangyu Xi, Shikun Zhang |
CIKM | 6 |
| 2021 | Cross-document Event Identity via Dense AnnotationabstractIn this paper, we study the identity of textual events from different documents.While the complex nature of event identity is previously studied (Hovy et al., 2013), the case of events across documents is unclear.Prior work on cross-document event coreference has two main drawbacks.First, they restrict the annotations to a limited set of event types.Second, they insufficiently tackle the concept of event identity.Such annotation setup reduces the pool of event mentions and prevents one from considering the possibility of quasiidentity relations.We propose a dense annotation approach for cross-document event coreference, comprising a rich source of event mentions and a dense annotation effort between related document pairs.To this end, we design a new annotation workflow with careful quality control and an easy-to-use annotation interface.In addition to the links, we further collect overlapping event contexts, including time, location, and participants, to shed some light on the relation between identity decisions and context.We present an open-access dataset for cross-document event coreference, CDEC-WN, collected from English Wikinews and open-source our annotation toolkit to encourage further research on cross-document tasks. 1 Adithya Pratapa, Zhengzhong Liu 0001, Kimihiro Hasegawa, Yukari Yamakawa, Shikun Zhang, Teruko Mitamura |
CoNLL | 6 |
| 2021 | Keyword-Aware Encoder for Abstractive Text Summarization
Tianxiang Hu, Jingxi Liang, Wei Ye 0004, Shikun Zhang |
DASFAA (2) | 4 |
| 2021 | Improving Event Detection by Exploiting Label HierarchyabstractEvent types are hierarchical, yet most existing methods for event detection classify candidate triggers into fine-grained event types directly, without considering the rich semantic correlations in the hierarchy of event types. To fully utilize such information to improve the detection of fine-grained event types, we propose a three-layer label hierarchy and introduce the detection of two coarser-grained types as auxiliary classification tasks. In particular, we leverage the supplementary supervision information from label hierarchy by a novel Logits Mapping (LM) strategy, which generates logits (the intermediate representations fed into classifier) for coarser-grained types by heuristic mapping of logits for fine-grained types. In this way, training signals provided by auxiliary tasks can help the encoder produce more precise logits via back propagation, thus providing a simple (no extra parameter needed) yet effective way to improve the target task. Results of extensive experiments on the ACE 2005 show that LM can not only be easily integrated into the state-of-the-art methods and achieve significant improvement over them, but also can effectively alleviate the data sparseness problem. Xiangyu Xi, Wei Ye 0004, Tong Zhang 0001, Quanxiu Wang, Shikun Zhang, Huixing Jiang, Wei Wu 0014 |
ICASSP | 5 |
| 2021 | Exploiting Method Names to Improve Code Summarization: A Deliberation Multi-Task Learning ApproachabstractCode summaries are brief natural language descriptions of source code pieces. The main purpose of code summarization is to assist developers in understanding code and to reduce documentation workload. In this paper, we design a novel multi-task learning (MTL) approach for code summarization through mining the relationship between method code summaries and method names. More specifically, since a method's name can be considered as a shorter version of its code summary, we first introduce the tasks of generation and informativeness prediction of method names as two auxiliary training objectives for code summarization. A novel two-pass deliberation mechanism is then incorporated into our MTL architecture to generate more consistent intermediate states fed into a summary decoder, especially when informative method names do not exist. To evaluate our deliberation MTL approach, we carried out a large-scale experiment on two existing datasets for Java and Python. The experiment results show that our technique can be easily applied to many state-of-the-art neural models for code summarization and improve their performance. Meanwhile, our approach shows significant superiority when generating summaries for methods with non-informative names. Rui Xie 0003, Wei Ye 0004, Jinan Sun, Shikun Zhang |
ICPC | 4 |
| 2021 | QuadrupletBERT: An Efficient Model For Embedding-Based Large-Scale RetrievalabstractPeiyang Liu, Sen Wang, Xi Wang, Wei Ye, Shikun Zhang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Peiyang Liu, Wei Ye 0004, Shikun Zhang |
NAACL-HLT | 5 |
| 2021 | Multi-Hop Transformer for Document-Level Machine TranslationabstractLong Zhang, Tong Zhang, Haibo Zhang, Baosong Yang, Wei Ye, Shikun Zhang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Long Zhang 0012, Tong Zhang 0001, Haibo Zhang 0013, Baosong Yang, Wei Ye 0004, Shikun Zhang |
NAACL-HLT | 6 |
| 2021 | Legal Judgment Prediction with Multi-Stage Case Representation Learning in the Real Court SettingabstractLegal judgment prediction(LJP) is an essential task for legal AI. While prior methods studied on this topic in a pseudo setting by employing the judge-summarized case narrative as the input to predict the judgment, neglecting critical case life-cycle information in real court setting could threaten the case logic representation quality and prediction correctness. In this paper, we introduce a novel challenging dataset from real courtrooms to predict the legal judgment in a reasonably encyclopedic manner by leveraging the genuine input of the case - plaintiff's claims and court debate data, from which the case's facts are automatically recognized by comprehensively understanding the multi-role dialogues of the court debate, and then learnt to discriminate the claims so as to reach the final judgment through multi-task learning. An extensive set of experiments with a large civil trial data set shows that the proposed model can more accurately characterize the interactions among claims, fact and debate for legal judgment prediction, achieving significant improvements over strong state-of-the-art baselines. Moreover, the user study conducted with real judges and law school students shows the neural predictions can also be interpretable and easily observed, and thus enhancing the trial efficiency and judgment quality. Xiaozhong Liu 0001, Wei Ye 0004, Changlong Sun, Shikun Zhang |
SIGIR | 7 |
| 2021 | "Did you know this camera tracks your mood?": Understanding Privacy Expectations and Preferences in the Age of Video AnalyticsabstractAbstract Cameras are everywhere, and are increasingly coupled with video analytics software that can identify our face, track our mood, recognize what we are doing, and more. We present the results of a 10-day in-situ study designed to understand how people feel about these capabilities, looking both at the extent to which they expect to encounter them as part of their everyday activities and at how comfortable they are with the presence of such technologies across a range of realistic scenarios. Results indicate that while some widespread deployments are expected by many (e.g., surveillance in public spaces), others are not, with some making people feel particularly uncomfortable. Our results further show that individuals’ privacy preferences and expectations are complicated and vary with a number of factors such as the purpose for which footage is captured and analyzed, the particular venue where it is captured, and whom it is shared with. Finally, we discuss the implications of people’s rich and diverse preferences on opt-in or opt-out rights for the collection and use (including sharing) of data associated with these video analytics scenarios as mandated by regulations. Because of the user burden associated with the large number of privacy decisions people could be faced with, we discuss how new types of privacy assistants could possibly be configured to help people manage these decisions. Shikun Zhang, Yuanyuan Feng, Lujo Bauer, Lorrie Faith Cranor, Anupam Das 0001, Norman M. Sadeh |
Proc. Priv. Enhancing Technol. | 1 |
| 2020 | Deep Dynamic Boosted ForestabstractRandom forest is widely exploited as an ensemble learning method. In many practical applications, however, there is still a significant challenge to learn from imbalanced data. To alleviate this limitation, we propose a deep dynamic boosted forest (DDBF), a novel ensemble algorithm that incorporates the notion of hard example mining into random forest. Specifically, we propose to measure the quality of each leaf node of every decision tree in the random forest to determine hard examples. By iteratively training and then removing easy examples from training data, we evolve the random forest to focus on hard examples dynamically so as to balance the proportion of samples and learn decision boundaries better. Data can be cascaded through these random forests learned in each iteration in sequence to generate more accurate predictions. Our DDBF outperforms random forest on 5 UCI datasets, MNIST and SATIMAGE, and achieved state-of-the-art results compared to other deep models. Moreover, we show that DDBF is also a new way of sampling and can be very useful and efficient when learning from imbalanced data. Haixin Wang 0003, Xingzhang Ren, Jinan Sun, Wei Ye 0004, Muzhi Yu, Shikun Zhang |
ACML | 7 |
| 2020 | Graph Enhanced Dual Attention Network for Document-Level Relation ExtractionabstractDocument-level relation extraction requires inter-sentence reasoning capabilities to capture local and global contextual information for multiple relational facts.To improve inter-sentence reasoning, we propose to characterize the complex interaction between sentences and potential relation instances via a Graph Enhanced Dual Attention network (GEDA).In GEDA, sentence representation generated by the sentence-to-relation (S2R) attention is refined and synthesized by a Heterogeneous Graph Convolutional Network before being fed into the relation-to-sentence (R2S) attention .We further design a simple yet effective regularizer based on the natural duality of the S2R and R2S attention, whose weights are also supervised by the supporting evidence of relation instances during training.An extensive set of experiments on an existing large-scale dataset show that our model achieves competitive performance, especially for the inter-sentence relation extraction, while the neural predictions can also be interpretable and easily observed. Bo Li 0099, Wei Ye 0004, Zhonghao Sheng, Rui Xie 0003, Xiangyu Xi, Shikun Zhang |
COLING | 6 |
| 2020 | Leveraging Human Prior Knowledge to Learn Sense Representations
Tong Zhang 0001, Wei Ye 0004, Xiangyu Xi, Long Zhang 0012, Shikun Zhang |
ECAI | 5 |
| 2020 | Stacking Networks Dynamically for Image Restoration Based on the Plug-and-Play Framework
Haixin Wang 0003, Muzhi Yu, Jinan Sun, Wei Ye 0004, Chen Wang 0005, Shikun Zhang |
ECCV (13) | 7 |
| 2020 | Sliding Hierarchical Recurrent Neural Networks for Sequence ClassificationabstractHierarchical Recurrent Neural Networks (HRNN) is an important advance in improving efficiency and performance of sequence classification in recent years. The intuition behind this approach is to slice long sequences into many short sub-sequences and process them in parallel, then capturing the long-term dependencies between those sub-sequences by deeper layers of the networks. In this paper, we propose a novel architecture called Sliding Hierarchical Recurrent Neural Network (SHRNN). We introduce a new sliding mechanism on the input sequence of each layer, named recursive block, so that SHRNN can process the input sequence effectively. We also introduce layer-wise attention and multi-layer regularization for further improvements. We perform large-scale experiments in sequence classification task of both text and image on 8 datasets. As result, we not only achieve new start-of-the-art performance on all datasets by SHRNN, but also investigate effects of different components of SHRNN systematically and thoroughly, which provides best practice for the usage of SHRNN. Bo Li 0099, Zhonghao Sheng, Wei Ye 0004, Shikun Zhang |
IJCNN | 6 |
| 2020 | Not All Synonyms Are Created Equal: Incorporating Similarity of Synonyms to Enhance Word EmbeddingsabstractTraditional word embedding approaches learn semantic information from the associated contexts of words on large unlabeled corpora, which ignores a fact that synonymy between words happens often within different contexts in a corpus, so this relationship will not be well embedded into vectors. Furthermore, existing synonymy-based models directly incorporate synonyms to train word embeddings, but still neglect the similarity between words and corresponding synonyms. In this paper, we explore a novel approach that employs the similarity between words and corresponding synonyms to train and enhance word embeddings. To this purpose, we build two Synonymy Similarity Models (SSMs), named SSM-W and SSM-M respectively, which adopt different strategies to incorporate the similarity between words and corresponding synonyms during the training process. We evaluated our models for both Chinese and English. The results demonstrate that our models outperform the baselines on seven word similarity datasets. For the analogical reasoning and text classification tasks, our models also surpass all the baselines including a synonymy-based model. Peiyang Liu, Wei Ye 0004, Xiangyu Xi, Shikun Zhang |
IJCNN | 6 |
| 2020 | Exploiting Code Knowledge Graph for Bug Localization via Bi-directional AttentionabstractBug localization automatic localize relevant source files given a natural language description of bug within a software project. For a large project containing hundreds and thousands of source files, developers need cost lots of time to understand bug reports generated by quality assurance and localize these buggy source files. Traditional methods are heavily depending on the information retrieval technologies which rank the similarity between source files and bug reports in lexical level. Recently, deep learning based models are used to extract semantic information of code with significant improvements for bug localization. However, programming language is a highly structural and logical language, which contains various relations within and cross source files. Thus, we propose KGBugLocator to utilize knowledge graph embeddings to extract these interrelations of code, and a keywords supervised bi-directional attention mechanism regularize model with interactive information between source files and bug reports. With extensive experiments on four different projects, we prove our model can reach the new the-state-of-art(SOTA) for bug localization. Rui Xie 0003, Wei Ye 0004, Shikun Zhang |
ICPC | 5 |
| 2020 | Leveraging Code Generation to Improve Code Retrieval and Summarization via Dual LearningabstractCode summarization generates brief natural language description given a source code snippet, while code retrieval fetches relevant source code given a natural language query. Since both tasks aim to model the association between natural language and programming language, recent studies have combined these two tasks to improve their performance. However, researchers have yet been able to effectively leverage the intrinsic connection between the two tasks as they train these tasks in a separate or pipeline manner, which means their performance can not be well balanced. In this paper, we propose a novel end-to-end model for the two tasks by introducing an additional code generation task. More specifically, we explicitly exploit the probabilistic correlation between code summarization and code generation with dual learning, and utilize the two encoders for code summarization and code generation to train the code retrieval task via multi-task learning. We have carried out extensive experiments on an existing dataset of SQL and Python, and results show that our model can significantly improve the results of the code retrieval task over the-state-of-art models, as well as achieve competitive performance in terms of BLEU score for the code summarization task. Wei Ye 0004, Rui Xie 0003, Tianxiang Hu, Xiaoyin Wang, Shikun Zhang |
WWW | 6 |
| 2020 | ADHD fMRI short-time analysis method for edge computing based on multi-instance learning
Chengfeng Dou, Shikun Zhang, Hanping Wang, Yu Huang 0004, Weihua Yue |
J. Syst. Archit. | 2 |
| 2020 | The Best of Both Worlds: Mitigating Trade-offs Between Accuracy and User Burden in Capturing Mobile App Privacy PreferencesabstractAbstract In today’s data-centric economy, data flows are increasingly diverse and complex. This is best exemplified by mobile apps, which are given access to an increasing number of sensitive APIs. Mobile operating systems have attempted to balance the introduction of sensitive APIs with a growing collection of permission settings, which users can grant or deny. The challenge is that the number of settings has become unmanageable. Yet research also shows that existing settings continue to fall short when it comes to accurately capturing people’s privacy preferences. An example is the inability to control mobile app permissions based on the purpose for which an app is requesting access to sensitive data. In short, while users are already overwhelmed, accurately capturing their privacy preferences would require the introduction of an even greater number of settings. A promising approach to mitigating this trade-off lies in using machine learning to generate setting recommendations or bundle some settings. This article is the first of its kind to offer a quantitative assessment of how machine learning can help mitigate this trade-off, focusing on mobile app permissions. Results suggest that it is indeed possible to more accurately capture people’s privacy preferences while also reducing user burden. Daniel Smullen, Yuanyuan Feng, Shikun Zhang, Norman M. Sadeh |
Proc. Priv. Enhancing Technol. | 3 |
| 2019 | Exploiting Entity BIO Tag Embeddings and Multi-task Learning for Relation Extraction with Imbalanced DataabstractIn practical scenario, relation extraction needs to first identify entity pairs that have relation and then assign a correct relation class.However, the number of non-relation entity pairs in context (negative instances) usually far exceeds the others (positive instances), which negatively affects a model's performance.To mitigate this problem, we propose a multitask architecture which jointly trains a model to perform relation identification with crossentropy loss and relation classification with ranking loss.Meanwhile, we observe that a sentence may have multiple entities and relation mentions, and the patterns in which the entities appear in a sentence may contain useful semantic information that can be utilized to distinguish between positive and negative instances.Thus we further incorporate the embeddings of character-wise/word-wise BIO tag from the named entity recognition task into character/word embeddings to enrich the input representation.Experiment results show that our proposed approach can significantly improve the performance of a baseline model with more than 10% absolute increase in F1-score, and outperform the state-of-theart models on ACE 2005 Chinese and English corpus.Moreover, BIO tag embeddings are particularly effective and can be used to improve other models as well.* indicates equal contribution. Wei Ye 0004, Bo Li 0099, Rui Xie 0003, Zhonghao Sheng, Shikun Zhang |
ACL (1) | 6 |
| 2019 | Capturing source code semantics via tree-based convolution over API-enhanced ASTabstractWhen deep learning meets big code, a key question is how to efficiently learn a distributed representation for source code that can capture its semantics effectively. We propose to use tree-based convolution over API-enhanced AST. To demonstrate the effectiveness of our approach, we apply it to detect semantic clones---code fragments with similar semantics but dissimilar syntax. Experiment results show that our approach outperforms an existing state-of-the-art approach that uses tree-based LSTM, with an increase of 0.39 and 0.12 in F1-score on OJClone and BigCloneBench respectively. We further propose architectures that incorporate our approach for code search and code summarization. Wei Ye 0004, Shikun Zhang |
CF | 3 |
| 2019 | Long Text Analysis Using Sliced Recurrent Neural Networks with Breaking Point Information EnrichmentabstractSliced recurrent neural networks (SRNNs) are the state-of-the-art efficient solution for long text analysis tasks; however, their slicing operations inevitably result in long-term dependency loss in lower-level networks and thus limit their accuracy. Therefore, we propose a breaking point information enrichment mechanism to strengthen dependencies between sliced subsequences without hindering parallelization. Then, the resulting BPIE-SRNN model is further extended to a bidirectional model, BPIE-BiSRNN, to utilize the dependency information in not only the previous but also the following contexts. Experiments on four large public real-world datasets demonstrate that the BPIE-SRNN and BPIE-BiSRNN models always achieve a much better accuracy than SRNNs and BiSRNNs, while maintaining a superior training efficiency. Bo Li 0099, Zehua Cheng, Zhenghua Xu 0001, Wei Ye 0004, Thomas Lukasiewicz, Shikun Zhang |
ICASSP | 6 |
| 2019 | A Hybrid Character Representation for Chinese Event DetectionabstractFor the Chinese language, event triggers in a sentence may appear inside or across words after word segmentation. Thus recent works on Chinese event detection often formulate the task as a character-wise sequence labeling problem instead of a word-wise one. Due to a limited amount of corpus, however, it is more difficult in practice to train character-wise models to capture the inner structure of event triggers and the semantics of sentence-level context compared with word-wise ones. In this paper, we propose to improve character-wise models by incorporating word information and language model representation into Chinese character representation. More specifically, the former consists of the position of the character inside a word and the word's embedding, which can aid structural pattern learning; the latter is obtained by BERT, which contains long-distance semantic information. We construct a sequence tagging model equipped with the hybrid representation and evaluate our model on ACE 2005 Chinese corpus. Experiment results show that both word information and language model representation are effective enhancements, and our model gains an increase of 4.5 (6.5%) and 6.1 (9.4%) in F1-score in event trigger identification task and classification task respectively over the state-of-the-art method. Xiangyu Xi, Tong Zhang 0001, Wei Ye 0004, Rui Xie 0003, Shikun Zhang |
IJCNN | 6 |
| 2019 | DeepLink: A Code Knowledge Graph Based Deep Learning Approach for Issue-Commit Link RecoveryabstractLinks between issue reports and corresponding code commits to fix them can greatly reduce the maintenance costs of a software project. More often than not, however, these links are missing and thus cannot be fully utilized by developers. Current practices in issue-commit link recovery extract text features and code features in terms of textual similarity from issue reports and commit logs to train their models. These approaches are limited since semantic information could be lost. Furthermore, few of them consider the effect of source code files related to a commit on issue-commit link recovery, let alone the semantics of code context. To tackle these problems, we propose to construct code knowledge graph of a code repository and generate embeddings of source code files to capture the semantics of code context. We also use embeddings to capture the semantics of issue- or commit-related text. Then we use these embeddings to calculate semantic similarity and code similarity using a deep learning approach before training a SVM binary classification model with additional features. Evaluations on real-world projects show that our approach DeepLink can outperform the state-of-the-art method. Rui Xie 0003, Wei Ye 0004, Tianxiang Hu, Dongdong Du, Shikun Zhang |
SANER | 7 |
| 2018 | Runtime Resource Management for Microservices-Based Applications: A Congestion Game Approach (Short Paper)
Ruici Luo, Wei Ye 0004, Jinan Sun, Xueyang Liu, Shikun Zhang |
CollaborateCom | 5 |
| 2018 | Attention Enhanced Chinese Word Embeddings
Xingzhang Ren, Wei Ye 0004, Hang Hua, Shikun Zhang |
ICANN (1) | 5 |
| 2018 | Refining Traceability Links Between Vulnerability and Software Component in a Vulnerability Knowledge Graph
Dongdong Du, Xingzhang Ren, Jien Chen, Wei Ye 0004, Jinan Sun, Xiangyu Xi, Shikun Zhang |
ICWE | 9 |
| 2018 | CoBOT: static C/C++ bug detection in the presence of incomplete codeabstractTo obtain precise and sound results, most of existing static analyzers require whole program analysis with complete source code. However, in reality, the source code of an application always interacts with many third-party libraries, which are often not easily accessible to static analyzers. Worse still, more than 30% of legacy projects [1] cannot be compiled easily due to complicated configuration environments (e.g., third-party libraries, compiler options and macros), making ideal "whole-program analysis" unavailable in practice. This paper presents CoBOT [2], a static analysis tool that can detect bugs in the presence of incomplete code. It analyzes function APIs unavailable in application code by either using function summarization or automatically downloading and analyzing the corresponding library code as inferred from the application code and its configuration files. The experiments show that CoBOT is not only easy to use, but also effective in detecting bugs in real-world programs with incomplete code. Our demonstration video is at: https://youtu.be/bhjJp3e7LPM. Sen Ma, Sihao Shao, Yulei Sui, Fuyao Duan, Shikun Zhang |
ICPC | 10 |
| 2018 | Method and System for Detecting Anomalous User Behaviors: An Ensemble ApproachabstractMalicious user behavior that does not trigger access violation or data leak alert is difficult to detect.Using the stolen login credentials, the intruder doing espionage will first try to stay undetected, silently collect data that he is authorized to access from the company network.This paper presents an overview of User Behavior Analytics Platform built to collect logs, extract features and detect anomalous users which may contain potential insider threats.Besides, a multi-algorithms ensemble, combining OCSVM, RNN and Isolation Forest, is introduced.The experiment showed that the system with an ensemble of unsupervised anomaly detection algorithms can detect abnormal user behavior patterns.The experiment results indicate that OCSVM and RNN suffer from anomalies in the training set, and iF orest gives more false positives and false negatives, while the ensemble of three algorithms has great performance and achieves recall 96.55% and accuracy 91.24% on average. Xiangyu Xi, Dongdong Du, Shikun Zhang |
SEKE | 7 |
| 2018 | An Ensemble Approach for Detecting Anomalous User BehaviorsabstractAn intruder of a company’s network may use stolen login credentials to silently collect sensitive data. Such malicious user behavior is difficult to detect as long as it does not trigger access violation or data leak alert. In this paper, we propose to use an ensemble of three unsupervised anomaly detection algorithms, namely OCSVM, RNN and Isolation Forest, to detect abnormal user behavior patterns. Besides, an User Behavior Analytics (UBA) Platform is proposed to collect logs, extract features and conduct experiments. The experiment results indicate that our algorithm outperforms each individual algorithm with recall of 96.55% and precision of 91.24% on average, while both OCSVM and RNN suffer from anomalies in the training set, and [Formula: see text] produces more false positives and false negatives in prediction. Xiangyu Xi, Tong Zhang 0001, Wei Ye 0004, Shikun Zhang, Dongdong Du |
Int. J. Softw. Eng. Knowl. Eng. | 5 |
| 2018 | Turtles, Locks, and Bathrooms: Understanding Mental Models of Privacy Through IllustrationabstractAbstract Are the many formal definitions and frameworks of privacy consistent with a layperson’s understanding of privacy? We explored this question and identified mental models and metaphors of privacy, conceptual tools that can be used to improve privacy tools, communication, and design for everyday users. Our investigation focused on a qualitative analysis of 366 drawings of privacy from laypeople, privacy experts, children, and adults. Illustrators all responded to the prompt “What does privacy mean to you?” We coded each image for content, identifying themes from established privacy frameworks and defining the visual and conceptual metaphors illustrators used to model privacy. We found that many non-expert drawings illustrated a strong divide between public and private physical spaces, while experts were more likely to draw nuanced data privacy spaces. Young children’s drawings focused on bedrooms, bathrooms, or cheating on schoolwork, and seldom addressed data privacy. The metaphors, themes, and symbols identified by these findings can be used for improving privacy communication, education, and design by inspiring and informing visual and conceptual strategies for reaching laypeople. Maggie Oates, Yama Ahmadullah, Abigail Marsh, Chelse Swoopes, Shikun Zhang, Rebecca Balebako, Lorrie Faith Cranor |
Proc. Priv. Enhancing Technol. | 5 |
| 2016 | Unsupervised Text Recap Extraction for TV Series
Shikun Zhang, Louis-Philippe Morency |
EMNLP | 2 |
| 2016 | Follow My Recommendations: A Personalized Privacy Assistant for Mobile App Permissions
Bin Liu 0017, Mads Schaarup Andersen, Florian Schaub, Hazim Almuhimedi, Shikun Zhang, Norman M. Sadeh, Yuvraj Agarwal, Alessandro Acquisti |
SOUPS | 5 |
| 2015 | Towards a Deployment System for Cloud ApplicationsabstractA sophisticated deployment system plays an important role in automating and improving the process of software delivery, especially for cloud applications.Since cloud applications usually consist of many components run on different virtual machines, i.e., EC2 instances, the deployment is time-consuming and error-prone, which may involves manual operations and complex scripts.We develop a deployment system aiming to accelerate cloud application delivery.First of all, we propose a component model and a connector model involving cloud feature.Then we present a component management system, in which component can be configured and instantiated rapidly based component inheritance and composition.Finally, we develop a novel deployment mechanism that can automate deployment process across multiple cloud instances.Experiment shows that our approach can reduce the build time and downtime so that it can speed up the delivery process of software application. Ruici Luo, Wei Ye 0004, Shikun Zhang |
SEKE | 3 |
| 2014 | Dual Subtitles as Parallel Corpora
Shikun Zhang, Wang Ling, Chris Dyer |
LREC | 1 |
| 2012 | BuOA: An Achitecture Style for Modular Web ApplicationsabstractThough Web development technologies have made a dramatic progress in past decades, Web applications are still with a monolithic architecture in terms of their deployment and mechanisms for resolving component interdependencies, imposing constraints on, e.g., partial and dynamic upgrade, distributed and parallel development. In this paper, we propose BuOA (Business unit Oriented Architecture), a novel architecture style for modular Web applications. Compared with traditional layered architecture styles, BuOA vertically decomposes Web applications into a group of BUs (Business Units) each of which implements a complete and cohesive business function. To establish loosely coupled relationship between BUs, interactions between them are categorized into four patterns: observing, injecting, weaving and binding. The paper first explores the BU model based on a three-dimensional view of Web applications, and then presents a connector model that abstracts the four interaction patterns. Practical toolkits and a framework for BuOA-based development are also introduced based on a concrete example. With BuOA, we can design and develop evolvable Web applications in a modular, parallel and collaborative way. Wei Ye 0004, Ruici Luo, Shikun Zhang, Xueyang Liu, Wenhui Hu |
APSEC | 3 |
| 2012 | A Data Collaboration Model for Collaborative Design Based on C-Net
Wenhui Hu, Wei Ye 0004, Shikun Zhang, Xuan Sun 0001 |
SEKE | 4 |
| 2011 | Data Uncertainty Model for Mashup
Wenhui Hu, Wei Ye 0004, Shikun Zhang |
SEKE | 4 |
| 2011 | A Novel Method for Formally Detecting RFID Event Using Petri Nets
Jinan Sun, Yu Huang 0004, Shikun Zhang, Chong-Yi Yuan |
SEKE | 4 |
| 2008 | Modeling and Analysis of WS-BPEL Business Processes Based on ServiceNetabstractWeb service composition involves the combination of a number of existing web services to create a value-added one. WS-BPEL is a promising language which describes Web service composition in form of business processes. However, WS-BPEL is an XML-based language and may suffer from ambiguities or some erroneous properties. The analysis and verification of business processes specified in WS-BPEL by a formal method has been a hot topic in the research community lately. In this paper, we propose a method to model and analyze WS-BPEL business processes based on ServiceNet, a special class of Petri nets. Unlike most of the existing work which analyze only control flow properties of WS-BPEL business processes but neglect data flow properties, our method models and analyzes both control properties and data properties of WS-BPEL business processes. We present the transformation rules of WS-BPEL business processes into ServiceNet and enrich the reduction rules of ServiceNet by applying them in some practical projects. Then the throughness of a business processes can be verified by reducing its ServiceNet representation based on some reduction rules. Moreover, we define some data-aspect properties of WS-BPEL business processes and give the corresponding checking algorithm. Haiqiang Dun, Yu Huang 0004, Shikun Zhang |
APSEC | 4 |
| 2007 | A Three-Layer Model for Business Processes - Process Logic, Case Semantics and Workflow Management
Chong-Yi Yuan, Shikun Zhang, Yu Huang 0004 |
J. Comput. Sci. Technol. | 3 |
| 2006 | A Workflow Process Mining Algorithm Based on Synchro-Net
Xing-Qi Huang, Shikun Zhang, Chong-Yi Yuan |
J. Comput. Sci. Technol. | 4 |
| 2002 | Hierarchical message bus-based software architectural styleabstractAs the size and complexity of software systems increase, the design and specification of overall system structure become more significant issues than the choice of algorithms and data structures of computation. An appropriate architecture for a system is a key element of its success. Based on the practice of Jadebird software production line, this paper proposes a software architectural style based on hierarchical message buses, named JB/HMB. In this style, the component model consists of external interfaces, static structure and dynamic behavior, which depicts a component from different aspects. Supported by message buses, components interact with one another by messages, which can be used to describe distributed and concurrent systems well. JB/HMB style supports stepwise decomposition and refinement, and runtime system evolution. Finally, characteristics of JB/HMB style are summarized as a conclusion, and future research directions are specified. Shikun Zhang, Fuqing Yang |
Sci. China Ser. F Inf. Sci. | 1 |