VLDB 2026 Research / reviewers in the wild / expert
Yucheng Zhou 0001
dblp:74/9454-1
· DBLP profile ↗
31ranked-venue papers
13as first author
31since 2021 · last 2026
0009-0006-9883-5621ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 10 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 12 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sim4Seg: Boosting Multimodal Multi-disease Medical Diagnosis Segmentation with Region-Aware Vision-Language Similarity MasksabstractDespite significant progress in pixel-level medical image analysis, existing medical image segmentation models rarely explore medical segmentation and diagnosis tasks jointly. However, it is crucial for patients that models can provide explainable diagnoses along with medical segmentation results. In this paper, we introduce a medical vision-language task named Medical Diagnosis Segmentation (MDS), which aims to understand clinical queries for medical images and generate the corresponding segmentation masks as well as diagnostic results. To facilitate this task, we first present the Multimodal Multi-disease Medical Diagnosis Segmentation (M3DS) dataset, containing diverse multimodal multi-disease medical images paired with their corresponding segmentation masks and diagnosis chain-of-thought, created via an automated diagnosis chain-of-thought generation pipeline. Moreover, we propose Sim4Seg, a novel framework that improves the performance of diagnosis segmentation by taking advantage of the Region-Aware Vision-Language Similarity to Mask (RVLS2M) module. To improve overall performance, we investigate a test-time scaling strategy for MDS tasks. Experimental results demonstrate that our method outperforms baselines in both segmentation and diagnosis. Lingran Song, Yucheng Zhou 0001, Jianbing Shen |
AAAI | 2 |
| 2026 | Less Is More: Vision Representation Compression for Efficient Video Generation with Large Language ModelsabstractVideo generation using Large Language Models (LLMs) has shown promising potential, effectively leveraging the extensive LLM infrastructure to provide a unified framework for multimodal understanding and content generation. However, these methods face critical challenges, i.e., token redundancy and inefficiencies arising from long sequences, which constrain their performance and efficiency compared to diffusion-based approaches. In this study, we investigate the impact of token redundancy in LLM-based video generation by information-theoretic analysis and propose Vision Representation Compression (VRC), a novel framework designed to achieve more in both performance and efficiency with less video token representations. VRC introduces learnable representation compressor and decompressor to compress video token representations, enabling autoregressive next-sequence prediction in a compact latent space. Our approach reduces redundancy, shortens token sequences, and improves model's ability to capture underlying video structures. Our experiments demonstrate that VRC reduces token sequence lengths by a factor of 4, achieving more than 9~14x acceleration in inference while maintaining performance comparable to state-of-the-art video generation models. VRC not only accelerates the inference but also significantly reduces memory requirements during both model training and inference. Yucheng Zhou 0001, Jihai Zhang 0002, Guanjie Chen, Jianbing Shen, Yu Cheng 0001 |
AAAI | 1 |
| 2026 | Multimodal Large Language Models for Multi-Subject In-Context Image GenerationabstractRecent advances in text-to-image (T2I) generation have enabled visually coherent image synthesis from descriptions, but generating images containing multiple given subjects remains challenging.As the number of reference identities increases, existing methods often suffer from subject missing and semantic drift.To address this problem, we propose MU-SIC, the first MLLM specifically designed for MUlti-Subject In-Context image generation.To overcome the data scarcity, we introduce an automatic and scalable data generation pipeline that eliminates the need for manual annotation.Furthermore, we enhance the model's understanding of multi-subject semantic relationships through a vision chain-of-thought (CoT) mechanism, guiding step-by-step reasoning from subject images to semantics and generation.To mitigate identity entanglement and manage visual complexity, we develop a novel semantics-driven spatial layout planning method and demonstrate its test-time scalability.By incorporating complex subject images during training, we improve the model's capacity for chained reasoning.In addition, we curate MSIC, a new benchmark tailored for multi-subject in-context generation.Experimental results demonstrate that MUSIC significantly surpasses other methods in both multiand single-subject scenarios. Yucheng Zhou 0001, Dubing Chen, Jianbing Shen |
ACL (1) | 1 |
| 2026 | Compatibility-Aware Dynamic Fine-Tuning for Large Language ModelsabstractSupervised Fine-Tuning (SFT) is the predominant paradigm for aligning large language models (LLMs), yet it suffers from optimization instability and limited generalization.Recent work attributes this issue to pathological gradient scaling and proposes Dynamic Fine-Tuning (DFT) to correct it at the token level.However, DFT assumes all demonstrations are equally suitable learning targets, an assumption violated by the strong heterogeneity of large-scale instruction data, where demonstration-policy mismatch induces highvariance updates at the sample level.We introduce Compatibility-Aware Dynamic Fine-Tuning (CADFT), a principled extension of DFT that controls sample-level optimization variance.CADFT derives a dynamic, policydependent compatibility signal from model likelihoods to modulate supervised updates, suppressing high-variance gradients from incompatible demonstrations.We further propose a delayed, low-frequency compatibilityguided rewriting strategy to transform persistently incompatible demonstrations into learnable targets.We show that CADFT can be interpreted as a variance-controlled estimator that generalizes token-level stabilization in DFT to the sample level.Extensive experiments demonstrate improved stability, generalization, and cold-start reinforcement learning initialization, while remaining fully supervised and independent of explicit reward modeling. Yucheng Zhou 0001, Junwei Sheng, Qianning Wang, Jianbing Shen |
ACL (1) | 1 |
| 2026 | TheraMind: A Strategic and Adaptive Agent for Longitudinal Psychological Counseling
He Hu 0008, Chiyuan Ma, Qianning Wang, Lin Liu 0016, Yucheng Zhou 0001, Laizhong Cui, Fei Ma 0006, Qi Tian 0001 |
WWW | 5 |
| 2026 | Kardia-R1: Unleashing LLMs to Reason toward Understanding and Empathy for Emotional Support via Rubric-as-Judge Reinforcement Learning
Zhiqing Cui, Yuansheng Gao, Yucheng Zhou 0001, Usman Naseem |
WWW | 5 |
| 2026 | MindDialog: A large-scale benchmark for counseling dialogue understanding and generation
He Hu 0008, Juzheng Si, Qianning Wang, Tengjin Weng, Yihong Ji, Jiyue Jiang, Fei Ma 0006, Yucheng Zhou 0001, Laizhong Cui, Qi Tian 0001 |
Pattern Recognit. | 8 |
| 2025 | Improving Medical Large Vision-Language Models with Abnormal-Aware FeedbackabstractExisting Medical Large Vision-Language Models (Med-LVLMs), encapsulating extensive medical knowledge, demonstrate excellent capabilities in understanding medical images. However, there remain challenges in visual localization in medical images, which is crucial for abnormality detection and interpretation. To address these issues, we propose a novel UMed-LVLM designed to unveil medical abnormalities. Specifically, we collect a Medical Abnormalities Unveiling (MAU) dataset and propose a two-stage training method for UMed-LVLM training. To collect MAU dataset, we propose a prompt method utilizing the GPT-4V to generate diagnoses based on identified abnormal areas in medical images. Moreover, the two-stage training method includes Abnormal-Aware Instruction Tuning and Abnormal-Aware Rewarding, comprising Relevance Reward, Abnormal Localization Reward and Vision Relevance Reward. Experimental results demonstrate that our UMed-LVLM significantly outperforms existing Med-LVLMs in identifying and understanding medical abnormalities, achieving a 58% improvement over the baseline. In addition, this work shows that enhancing the abnormality detection capabilities of Med-LVLMs significantly improves their understanding of medical images and generalization capability. Our code and data release at URL. Yucheng Zhou 0001, Lingran Song, Jianbing Shen |
ACL (1) | 1 |
| 2025 | Safety Alignment via Constrained Knowledge UnlearningabstractZesheng Shi, Yucheng Zhou, Jing Li, Yuxin Jin, Yu Li, Daojing He, Fangming Liu, Saleh Alharbi, Jun Yu, Min Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zesheng Shi, Yucheng Zhou 0001, Jing Li 0034, Yu Li 0007, Daojing He, Fangming Liu, Saleh Alharbi, Jun Yu 0002, Min Zhang 0005 |
ACL (1) | 2 |
| 2025 | Impromptu Cybercrime Euphemism DetectionabstractDetecting euphemisms is essential for content security on various social media platforms, but existing methods designed for detecting euphemisms are ineffective in impromptu euphemisms. In this work, we make a first attempt to an exploration of impromptu euphemism detection and introduce the Impromptu Cybercrime Euphemisms Detection (ICED) dataset. Moreover, we propose a detection framework tailored to this problem, which employs context augmentation modeling and multi-round iterative training. Our detection framework mainly consists of a coarse-grained and a fine-grained classification model. The coarse-grained classification model removes most of the harmless content in the corpus to be detected. The fine-grained model, impromptu euphemisms detector, integrates context augmentation and multi-round iterations training to better predicts the actual meaning of a masked token. In addition, we leverage ChatGPT to evaluate the mode’s capability. Experimental results demonstrate that our approach achieves a remarkable 76-fold improvement compared to the previous state-of-the-art euphemism detector. Xiang Li 0001, Yucheng Zhou 0001, Laiping Zhao, Jing Li 0034, Fangming Liu |
COLING | 2 |
| 2025 | InsectMamba: State Space Model with Adaptive Composite Features for Insect RecognitionabstractThe recognition of insect pests is a critical task in agricultural technology, vital for ensuring food security and environmental sustainability. However, due to factors like high camouflage and species diversity, the complexity of pest identification poses significant obstacles. Existing methods struggle with fine-grained feature extraction to distinguish between closely related pest species. Although recent advancements have utilized modified network structures and combined deep learning approaches to improve accuracy, challenges persist due to the similarity between pests and their surroundings. To address this problem, we introduce InsectMamba, a novel approach that integrates State Space Models (SSMs), Convolutional Neural Networks (CNNs), Multi-Head Self-Attention mechanism (MSA), and Multilayer Perceptrons (MLPs) within Mix-SSM blocks. This integration facilitates the extraction of comprehensive visual features by leveraging the strengths of each encoding strategy. A selective module is also proposed to composite these features adaptively, enhancing the model’s ability to discern pest characteristics. InsectMamba was evaluated against strong competitors across insect classification and detection datasets. The results demonstrate its superior performance and verify the significance of each model component by an ablation study. Qianning Wang, Chenglin Wang 0010, Zhixin Lai, Yucheng Zhou 0001 |
ICASSP | 4 |
| 2025 | Semantic Causality-Aware Vision-Based 3D Occupancy PredictionabstractVision-based 3D semantic occupancy prediction is a critical task in 3D vision that integrates volumetric 3D reconstruction with semantic understanding. Existing methods, however, often rely on modular pipelines. These modules are typically optimized independently or use pre-configured inputs, leading to cascading errors. In this paper, we address this limitation by designing a novel causal loss that enables holistic, end-to-end supervision of the modular 2D-to-3D transformation pipeline. Grounded in the principle of 2D-to-3D semantic causality, this loss regulates the gradient flow from 3D voxel representations back to the 2D features. Consequently, it renders the entire pipeline differentiable, unifying the learning process and making previously non-trainable components fully learnable. Building on this principle, we propose the Semantic Causality-Aware 2D-to-3D Transformation, which comprises three components guided by our causal loss: Channel-Grouped Lifting for adaptive semantic mapping, Learnable Camera Offsets for enhanced robustness against camera perturbations, and Normalized Convolution for effective feature propagation. Extensive experiments demonstrate that our method achieves state-of-the-art performance on the Occ3D benchmark, demonstrating significant robustness to camera perturbations and improved 2D-to-3D semantic consistency. Dubing Chen, Yucheng Zhou 0001, Xianfei Li, Wenlong Liao, Jianbing Shen |
ICCV | 3 |
| 2025 | Towards Stabilized and Efficient Diffusion Transformers Through Long-Skip-Connections With Spectral Constraints
Guanjie Chen, Yucheng Zhou 0001, Xiaoye Qu, Tianlong Chen 0001, Yu Cheng 0001 |
ICCV | 3 |
| 2025 | DC-ControlNet: Decoupling Inter- and Intra-Element Conditions in Image Generation with Diffusion ModelsabstractIn this paper, we introduce DC (Decouple)-ControlNet, a highly flexible and precisely controllable framework for multi-condition image generation. The core idea behind DC-ControlNet is to decouple control conditions, transforming global control into a hierarchical system that integrates distinct elements, contents, and layouts. This enables users to mix these individual conditions with greater flexibility, leading to more efficient and accurate image generation control. Previous ControlNet-based models rely solely on global conditions, which affect the entire image and lack the ability of element- or region-specific control. This limitation reduces flexibility and can cause condition misunderstandings in multi-conditional image generation. To address these challenges, we propose both intra-element and Inter-element Controllers in DC-ControlNet. The Intra-Element Controller handles different types of control signals within individual elements, accurately describing the content and layout characteristics of the object. For interactions between elements, we introduce the Inter-Element Controller, which accurately handles multi-element interactions and occlusion based on user-defined relationships. Extensive evaluations show that DC-ControlNet significantly outperforms existing ControlNet models and Layout-to-Image generative models in terms of control flexibility and precision in multi-condition control. Our project website is available at: https://um-lab.github.io/DC-ControlNet/ Wencheng Han, Yucheng Zhou 0001, Jianbing Shen |
ICCV | 3 |
| 2025 | Weak to Strong Generalization for Large Language Models with Multi-capabilitiesabstractAs large language models (LLMs) grow in sophistication, some of their capabilities surpass human abilities, making it essential to ensure their alignment with human values and intentions, i.e., Superalignment. This superalignment challenge is particularly critical for complex tasks, as annotations provided by humans, as weak supervisors, may be overly simplistic, incomplete, or incorrect. Previous work has demonstrated the potential of training a strong model using the weak dataset generated by a weak model as weak supervision. However, these studies have been limited to a single capability. In this work, we conduct extensive experiments to investigate weak to strong generalization for LLMs with multi-capabilities. The experiments reveal that different capabilities tend to remain relatively independent in this generalization, and the effectiveness of weak supervision is significantly impacted by the quality and diversity of the weak datasets. Moreover, the self-bootstrapping of the strong model leads to performance degradation due to its overconfidence and the limited diversity of its generated dataset. To address these issues, we proposed a novel training framework using reward models to select valuable data, thereby providing weak supervision for strong model training. In addition, we propose a two-stage training method on both weak and selected datasets to train the strong model. Experimental results demonstrate our method significantly improves the weak to strong generalization with multi-capabilities. Yucheng Zhou 0001, Jianbing Shen, Yu Cheng 0001 |
ICLR | 1 |
| 2025 | ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional DependenciesabstractText-driven image editing has achieved remarkable success in following single instructions. However, real-world scenarios often involve complex, multi-step instructions, particularly ''chain'' instructions where operations are interdependent. Current models struggle with these intricate directives, and existing benchmarks inadequately evaluate such capabilities. Specifically, they often overlook multi-instruction and chain-instruction complexities, and common consistency metrics are flawed. To address this, we introduce ComplexBench-Edit, a novel benchmark designed to systematically assess model performance on complex, multi-instruction, and chain-dependent image editing tasks. ComplexBench-Edit also features a new vision consistency evaluation method that accurately assesses non-modified regions by excluding edited areas. Furthermore, we propose a simple yet powerful Chain-of-Thought (CoT)-based approach that significantly enhances the ability of existing models to follow complex instructions. Our extensive experiments demonstrate ComplexBench-Edit's efficacy in differentiating model capabilities and highlight the superior performance of our CoT-based method in handling complex edits. The data and code are released at https://github.com/llllly26/ComplexBench-Edit. Chenglin Wang 0010, Yucheng Zhou 0001, Qianning Wang, Kai Zhang 0001 |
ACM Multimedia | 2 |
| 2025 | Alternate Geometric and Semantic Denoising Diffusion for Protein Inverse Folding
Chenglin Wang 0010, Yucheng Zhou 0001, Zijie Zhai, Jianbing Shen, Kai Zhang 0001 |
ECML/PKDD (3) | 2 |
| 2024 | Fine-Grained Distillation for Long Document RetrievalabstractLong document retrieval aims to fetch query-relevant documents from a large-scale collection, where knowledge distillation has become de facto to improve a retriever by mimicking a heterogeneous yet powerful cross-encoder. However, in contrast to passages or sentences, retrieval on long documents suffers from the \textit{scope hypothesis} that a long document may cover multiple topics. This maximizes their structure heterogeneity and poses a granular-mismatch issue, leading to an inferior distillation efficacy. In this work, we propose a new learning framework, fine-grained distillation (FGD), for long-document retrievers. While preserving the conventional dense retrieval paradigm, it first produces global-consistent representations crossing different fine granularity and then applies multi-granular aligned distillation merely during training. In experiments, we evaluate our framework on two long-document retrieval benchmarks, which show state-of-the-art performance. Yucheng Zhou 0001, Tao Shen 0001, Xiubo Geng, Chongyang Tao, Jianbing Shen, Guodong Long, Can Xu 0002, Daxin Jiang |
AAAI | 1 |
| 2024 | SURf: Teaching Large Vision-Language Models to Selectively Utilize Retrieved InformationabstractLarge Vision-Language Models (LVLMs) have become pivotal at the intersection of computer vision and natural language processing.However, the full potential of LVLMs' Retrieval-Augmented Generation (RAG) capabilities remains underutilized.Existing works either focus solely on the text modality or are limited to specific tasks.Moreover, most LVLMs struggle to selectively utilize retrieved information and are sensitive to irrelevant or misleading references.To address these challenges, we propose a self-refinement framework designed to teach LVLMs to Selectively Utilize Retrieved Information (SURf).Specifically, when given questions that are incorrectly answered by the LVLM backbone, we obtain references that help correct the answers (positive references) and those that do not (negative references).We then fine-tune the LVLM backbone using a combination of these positive and negative references.Our experiments across three tasks and seven datasets demonstrate that our framework significantly enhances LVLMs' ability to effectively utilize retrieved multimodal references and improves their robustness against irrelevant or misleading information.The source code is available at https://github.com/GasolSun36/SURf. * Work done during internship at Shanghai AI Laboratory.† Both are corresponding authors.How many apples in the images? VQAVanilla: There are three apples.The image depicting five apples on a tree...The picture shows 7 apples .... leaves...Ours: There are four apples.Describe this image in details. Captioning Vanilla: A person walking in snow.The image depicting a...the skier is in a crouched position... The image captures a dynamic scene ..a skier dressed in a ... Jiashuo Sun, Jihai Zhang 0002, Yucheng Zhou 0001, Zhaochen Su, Xiaoye Qu, Yu Cheng 0001 |
EMNLP | 3 |
| 2024 | Multi-Modal Inductive Framework for Text-Video Retrieval
Qian Li 0033, Yucheng Zhou 0001, Cheng Ji 0001, Feihong Lu, Jianian Gong, Shangguang Wang, Jianxin Li 0002 |
ACM Multimedia | 2 |
| 2024 | MAGIS: LLM-Based Multi-Agent Framework for GitHub Issue ResolutionabstractIn software development, resolving the emergent issues within GitHub repositories is a complex challenge that involves not only the incorporation of new code but also the maintenance of existing code.
Large Language Models (LLMs) have shown promise in code generation but face difficulties in resolving Github issues, particularly at the repository level.
To overcome this challenge, we empirically study the reason why LLMs fail to resolve GitHub issues and analyze the major factors.
Motivated by the empirical findings, we propose a novel LLM-based **M**ulti-**A**gent framework for **G**itHub **I**ssue re**S**olution, **MAGIS**, consisting of four agents customized for software evolution: Manager, Repository Custodian, Developer, and Quality Assurance Engineer agents.
This framework leverages the collaboration of various agents in the planning and coding process to unlock the potential of LLMs to resolve GitHub issues.
In experiments, we employ the SWE-bench benchmark to compare MAGIS with popular LLMs, including GPT-3.5, GPT-4, and Claude-2.
MAGIS can resolve **13.94%** GitHub issues, significantly outperforming the baselines.
Specifically, MAGIS achieves an eight-fold increase in resolved ratio over the direct application of GPT-4, the advanced LLM. Wei Tao 0003, Yucheng Zhou 0001, Yanlin Wang 0001, Hongyu Zhang 0002, Yu Cheng 0001 |
NeurIPS | 2 |
| 2024 | KADEL: Knowledge-Aware Denoising Learning for Commit Message GenerationabstractCommit messages are natural language descriptions of code changes, which are important for software evolution such as code understanding and maintenance. However, previous methods are trained on the entire dataset without considering the fact that a portion of commit messages adhere to good practice (i.e., good-practice commits), while the rest do not. On the basis of our empirical study, we discover that training on good-practice commits significantly contributes to the commit message generation. Motivated by this finding, we propose a novel knowledge-aware denoising learning method called KADEL. Considering that good-practice commits constitute only a small proportion of the dataset, we align the remaining training samples with these good-practice commits. To achieve this, we propose a model that learns the commit knowledge by training on good-practice commits. This knowledge model enables supplementing more information for training samples that do not conform to good practice. However, since the supplementary information may contain noise or prediction errors, we propose a dynamic denoising training method. This method composes a distribution-aware confidence function and a dynamic distribution list, which enhances the effectiveness of the training process. Experimental results on the whole MCMD dataset demonstrate that our method overall achieves state-of-the-art performance compared with previous methods. Wei Tao 0003, Yucheng Zhou 0001, Yanlin Wang 0001, Hongyu Zhang 0002, Haofen Wang |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2023 | Multimodal Event Transformer for Image-guided Story Ending GenerationabstractImage-guided story ending generation (IgSEG) is to generate a story ending based on given story plots and ending image.Existing methods focus on cross-modal feature fusion but overlook reasoning and mining implicit information from story plots and ending image.To tackle this drawback, we propose a multimodal event transformer, an event-based reasoning framework for IgSEG.Specifically, we construct visual and semantic event graphs from story plots and ending image, and leverage event-based reasoning to reason and mine implicit information in a single modality.Next, we connect visual and semantic event graphs and utilize cross-modal fusion to integrate differentmodality features.In addition, we propose a multimodal injector to adaptive pass essential information to decoder.Besides, we present an incoherence detection to enhance the understanding context of a story plot and the robustness of graph modeling for our model.Experimental results show that our method achieves state-of-the-art performance for the image-guided story ending generation. Yucheng Zhou 0001, Guodong Long |
EACL | 1 |
| 2023 | Improving Cross-modal Alignment for Text-Guided Image InpaintingabstractText-guided image inpainting (TGII) aims to restore missing regions based on a given text in a damaged image.Existing methods are based on a strong vision encoder and a crossmodal fusion model to integrate cross-modal features.However, these methods allocate most of the computation to visual encoding, while light computation on modeling modality interactions.Moreover, they take cross-modal fusion for depth features, which ignores a finegrained alignment between text and image.Recently, vision-language pre-trained models (VLPM), encapsulating rich cross-modal alignment knowledge, have advanced in most multimodal tasks.In this work, we propose a novel model for TGII by improving cross-modal alignment (CMA).CMA model consists of a VLPM as a vision-language encoder, an image generator and global-local discriminators.To explore cross-modal alignment knowledge for image restoration, we introduce cross-modal alignment distillation and in-sample distribution distillation.In addition, we employ adversarial training to enhance the model to fill the missing region in complicated structures effectively.Experiments are conducted on two popular vision-language datasets.Results show that our model achieves state-of-the-art performance compared with other strong competitors. Yucheng Zhou 0001, Guodong Long |
EACL | 1 |
| 2023 | Disentangled and Robust Representation Learning for Bragging Classification in Social MediaabstractResearching bragging behavior on social media arouses interest of computational (socio) linguists. However, existing bragging classification datasets suffer from a serious data imbalance issue. Because labeling a data-balance dataset is expensive, most methods introduce external knowledge to improve model learning. Nevertheless, such methods inevitably introduce noise and non-relevance information from external knowledge. To overcome the drawback, we propose a novel bragging classification method with disentangle-based representation augmentation and domain-aware adversarial strategy. Specifically, model learns to disentangle and reconstruct representation and generate augmented features via disentangle-based representation augmentation. Moreover, domain-aware adversarial strategy aims to constrain domain of augmented features to improve their robustness. Experimental results demonstrate that our method achieves state-of-the-art performance compared to other methods. Xiang Li 0001, Yucheng Zhou 0001 |
ICASSP | 2 |
| 2023 | Topic-Selective Graph Network for Topic-Focused Summarization
Zesheng Shi, Yucheng Zhou 0001 |
PAKDD (4) | 2 |
| 2022 | ClarET: Pre-training a Correlation-Aware Context-To-Event Transformer for Event-Centric Generation and ClassificationabstractGenerating new events given context with correlated ones plays a crucial role in many eventcentric reasoning tasks.Existing works either limit their scope to specific scenarios or overlook event-level correlations.In this paper, we propose to pre-train a general Correlationaware context-to-Event Transformer (ClarET) for event-centric reasoning.To achieve this, we propose three novel event-centric objectives, i.e., whole event recovering, contrastive eventcorrelation encoding and prompt-based event locating, which highlight event-level correlations with effective training.The proposed ClarET is applicable to a wide range of eventcentric reasoning scenarios, considering its versatility of (i) event-correlation types (e.g., causal, temporal, contrast), (ii) application formulations (i.e., generation and classification), and (iii) reasoning types (e.g., abductive, counterfactual and ending reasoning).Empirical fine-tuning results, as well as zero-and fewshot learning, on 9 benchmarks (5 generation and 4 classification tasks covering 4 reasoning types with diverse event correlations), verify its effectiveness and generalization ability. Yucheng Zhou 0001, Tao Shen 0001, Xiubo Geng, Guodong Long, Daxin Jiang |
ACL (1) | 1 |
| 2022 | Sketch StorytellingabstractSketch storytelling aims to generate a story for a given sketch. Although image captioning based on deep learning has great progress, describing the sketch in a story style is still a challenge. The reason is that there is currently no paired sketch-story data which is expensive to acquire. Therefore, it is necessary to train a sketch storytelling model without using any paired sketch-story data. To address these issues, we replace the natural image in image caption dataset with the sketch with the corresponding objects to generate pseudo sketch, which can obtain pseudo paired sketch-caption and sketch-image data. Due to these pseudo sketches are not drawn in a standardized way, we present a selective attention module to reduce noise for pseudo sketches. Furthermore, we propose four novel objectives include sketch-image matching, image-caption generation, sketch-caption generation, mask infilling, which help the model learn mappings between sketch and story from more perspectives. Consequently, we built a test set for sketch-story evaluation. The experimental results show that our model achieves state-of-the-art performance as compared to other methods. Yucheng Zhou 0001 |
ICASSP | 1 |
| 2022 | EventBERT: A Pre-Trained Model for Event Correlation ReasoningabstractEvent correlation reasoning infers whether a natural language paragraph containing multiple events conforms to human common sense. For example, “Andrew was very drowsy, so he took a long nap, and now he is very alert” is sound and reasonable. In contrast, “Andrew was very drowsy, so he stayed up a long time, now he is very alert” does not comply with human common sense. Such reasoning capability is essential for many downstream tasks, such as script reasoning, abductive reasoning, narrative incoherence, story cloze test, etc. However, conducting event correlation reasoning is challenging due to a lack of large amounts of diverse event-based knowledge and difficulty in capturing correlation among multiple events. In this paper, we propose EventBERT, a pre-trained model to encapsulate eventuality knowledge from unlabeled text. Specifically, we collect a large volume of training examples by identifying natural language paragraphs that describe multiple correlated events and further extracting event spans in an unsupervised manner. We then propose three novel event- and correlation-based learning objectives to pre-train an event correlation model on our created training corpus. Experimental results show EventBERT outperforms strong baselines on four downstream tasks, and achieves state-of-the-art results on most of them. Moreover, it outperforms existing pre-trained models by a large margin, e.g., 6.5 ∼ 23%, in zero-shot learning of these tasks. Yucheng Zhou 0001, Xiubo Geng, Tao Shen 0001, Guodong Long, Daxin Jiang |
WWW | 1 |
| 2021 | Triple Sequence Generative Adversarial Nets for Unsupervised Image CaptioningabstractLabelling image-sentence is expensive and some unsupervised image captioning methods show promising results on caption generation. However, the generated captions are not very relevant to images due to the excessive dependence on the corpus. In order to overcome that drawback, we focus on the correspondence between image and sentence to construct an image caption with better mapping relation. In this paper, we present a novel triple sequence generative adversarial net including an image generator, a discriminator, and a sentence generator. The image generator is used to generate the image regions for words. Meanwhile, the sentence corpus guides the sentence generator based on the generated image regions. The discriminator judges the relevance between the words in the sentence and the generated image regions. In the experiments, we use a large number of unpaired images and sentences to train our model on the unsupervised and unpaired setting. The experimental results demonstrate that our method achieves significant improvements as compared to all baselines. Yucheng Zhou 0001, Wei Tao 0003 |
ICASSP | 1 |
| 2021 | Improving Zero-Shot Cross-lingual Transfer for Multilingual Question Answering over Knowledge GraphabstractYucheng Zhou, Xiubo Geng, Tao Shen, Wenqiang Zhang, Daxin Jiang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Yucheng Zhou 0001, Xiubo Geng, Tao Shen 0001, Daxin Jiang |
NAACL-HLT | 1 |