VLDB 2026 Research / reviewers in the wild / expert
Yin Zhang 0006
dblp:91/3045-6
· DBLP profile ↗
91ranked-venue papers
2as first author
46since 2021 · last 2026
0000-0001-6986-4227ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 40 · 1 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 11 since 2021Databases, data management, data science and information retrieval · 19 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 3 since 2021Systems, architecture and hardware · 3 · 2 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VideoPro: Adaptive Program Reasoning for Long Video UnderstandingabstractChenglin Li, Feng Han, Yikun Wang, Ruilin Li, Shuai Dong, Haowen Hou, Haitao Li, Qianglong Chen, Feng Tao, Jingqi Tong, Yin Zhang, Jiaqi Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yikun Wang 0001, Haowen Hou, Qianglong Chen, Jingqi Tong, Yin Zhang 0006 |
ACL (1) | 11 |
| 2026 | Leveraging Outline-Optimized Generative Interactions and Critique for Self-Refining Outlines with Reinforcement LearningabstractHengwei Liu, Haoyuan Ma, Qingqing Lyu, Daoxin Zhang, Yao Hu, Yongliang Shen, Yin Zhang, Weiming Lu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Hengwei Liu, Qingqing Lyu, Daoxin Zhang, Yongliang Shen 0001, Yin Zhang 0006, Weiming Lu 0001 |
ACL (1) | 7 |
| 2026 | AutoTaskEval: Towards Domain-Specific and Fine-Grained Evaluation for LLMsabstractQingqing Lyu, Linjuan Wu, Yongliang Shen, Hengwei Liu, Hao Li, Shengpei Jiang, Yin Zhang, Weiming Lu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Qingqing Lyu, Linjuan Wu, Yongliang Shen 0001, Hengwei Liu, Shengpei Jiang, Yin Zhang 0006, Weiming Lu 0001 |
ACL (1) | 7 |
| 2026 | Vector Calligrapher: Generating Scalable Vector Graphics via Structured Linguistic SupervisionabstractGenerating SVG-based fonts requires Multimodal Large Language Models (MLLMs) to translate high-level linguistic intent into lowlevel, topologically constrained symbolic sequences.However, current approaches struggle with two fundamental misalignments: the semantic ambiguity of unstructured natural language for precise geometric control, and the inefficiency of generic text tokenizers, which fragment coordinate-dense SVG XML into excessively long sequences with low information density.In this work, we propose Vector Calligrapher, a system that treats SVG generation as a conditional language modeling task optimized for both semantic grounding and representational efficiency. Bo Zhou 0025, Xikang Chen, Yin Zhang 0006 |
ACL (1) | 4 |
| 2026 | SmartNS: Enabling Line-rate and Flexible Network Stack with SmartNICabstractAs the gap between network and CPU speeds rapidly increases, the CPU-centric network stack proves inadequate due to excessive CPU and memory overheads. Though hardware-offloaded network stacks alleviate these issues, they suffer from limited flexibility in both control and data planes. It seems promising to offload network stacks to Smart-NICs to provide high flexibility. However, naive offloading leads to low throughput due to the inherent architectural limitations of widespread off-path SmartNICs. Even simple operations on staged network traffic would overwhelm the limited SmartNIC memory bandwidth. To this end, we design SmartNS, a SmartNIC-centric network stack with software transport programmability and line-rate packet processing capabilities. To tackle the limitations of SmartNIC-induced challenges, we propose a header-only offloading TX path and an unlimited-working-set in-cache processing RX path to minimize memory traffic to fit the wimpy SmartNIC memory bandwidth. To fully utilize the SmartNIC computing resources, we propose a programmable offloading engine to enable cloud providers to offload customized tasks along with the network stack processing. We prototype SmartNS using the widespread Nvidia BlueField-3 SmartNIC, and implement RoCEv2 and Solar transport protocols by leveraging SmartNS's software programmability. SmartNS achieves 2.2× higher throughput than the microkernel-based baseline in block storage disaggregation and 1.3× higher throughput than the hardware-offloaded baseline in KVCache transfer. Xuzheng Chen, Jie Zhang 0081, Baolin Zhu, Xueying Zhu, Zhongqing Chen, Lingjun Zhu, Yin Zhang 0006, Yuanchao Shu, Peng Cheng 0001, Zeke Wang |
EuroSys | 10 |
| 2026 | Exp2RL: Enhancing LLM Agent Training with Expert Experiences
Yicheng Li 0005, Qianglong Chen, Zhirui Zhang, Yin Zhang 0006 |
KSEM (3) | 6 |
| 2026 | Multimodal fusion with LLM content via hierarchical progressive transformer for explainable fake news detection
Yin Zhang 0006 |
Inf. Process. Manag. | 4 |
| 2025 | Effective Two-Stage Knowledge Transfer for Multi-Entity Cross-Domain RecommendationabstractIn recent years, the recommendation content on e-commerce platforms has become increasingly rich -- a single user feed may contain multiple entities, such as selling products, short videos, and content posts. To deal with the multi-entity cross-domain recommendation problem, an intuitive solution is to adopt the shared-network-based architecture for joint training. The underlying idea is to transfer the knowledge from one type of entity (source entity) to another (target entity). However, different from the conventional same-entity cross-domain recommendation, multi-entity knowledge transfer encounters several important challenges: (1) data distributions of the source entity and target entity are naturally different, making the shared-network-based joint training susceptible to the negative transfer issue(2) the corresponding feature schema of each entity is not exactly aligned (e.g., price is an essential feature for selling product while missing for content posts), making the existing approaches no longer appropriate. Recent researchers have also experimented with the pre-training and fine-tuning paradigm. Again, they only take into account the scenarios with the same entity type and feature schemas, which is inappropriate in our case. To this end, we design a pre-training & fine-tuning based Multi-entity Knowledge Transfer framework called MKT. MKT utilizes a multi-entity pre-training module to extract transferable knowledge across different entities. In particular, a feature alignment module is first applied to scale and align different feature schemas. Afterward, a couple of knowledge extractors are employed to extract the common and independent knowledge. In the end, the extracted common knowledge is adopted for target entity model training. Through extensive offline and online experiments on public and industrial datasets, we demonstrated the superiority of MKT over multiple State-Of-The-Art methods. MKT has also been deployed for content post recommendations on our production system. Jianyu Guan, Zongming Yin, Leihui Chen, Yin Zhang 0006, Fei Huang 0002, Shuguang Han, Jufeng Chen |
KDD (2) | 5 |
| 2025 | ASTNet: Asynchronous Spatio-Temporal Network for Large-Scale Chemical Sensor ForecastingabstractThe chemical industry is faced with the urgent challenge of effectively harnessing the vast amounts of time-series data generated by thousands of sensors, which is essential for forecasting chemical states, achieving accurate real-time control of production processes. Traditional forecasting methods suffer from high computational latency and struggle with the complexity of spatiotemporal dependencies. As a result, modeling this data becomes challenging. This paper introduces a novel approach, referred to as ASTNet, designed to address these challenges. ASTNet integrates an asynchronous spatiotemporal modeling framework that combines temporal and spatial encoders, enabling concurrent learning of temporal and spatial dependencies while reducing computational latency. Additionally, it introduces a gated graph fusion mechanism that adaptively combines static (meta) and evolving (dynamic) sensor graphs, enhancing the handling of heterogeneous sensor data and spatial correlations. Extensive experiments on three real-world chemical sensor datasets demonstrate that ASTNet outperforms SOTA methods in terms of both prediction accuracy and computational efficiency, making ASTNet successfully deployed in chemical engineering industrial scenarios. Shihao Tu, Yang Yang 0009, Wenyue Ding, Yicheng Lu, Qingkai Ren, Yin Zhang 0006 |
KDD (2) | 7 |
| 2025 | IU4Rec: Interest Unit-Based Product Organization and Recommendation for E-Commerce PlatformabstractMost recommendation systems typically follow a product-based paradigm utilizing user-product interactions to identify the most engaging items for users. However, this product-based paradigm has notable drawbacks for Xianyu~. Xianyu is China's largest online C2C e-commerce platform where a large portion of the product are post by individual sellers. Most of the product on Xianyu posted from individual sellers often have limited stock available for distribution, and once the product is sold, it's no longer available for distribution. This result in most items distributed product on Xianyu having relatively few interactions, affecting the effectiveness of traditional recommendation depending on accumulating user-item interactions. To address these issues, we introduce IU4Rec, an Interest Unit-based two-stage Recommendation system framework. We first group products into clusters based on attributes such as category, image, and semantics. These IUs are then integrated into the Recommendation system, delivering both product and technological innovations. IU4Rec begins by grouping products into clusters based on attributes such as category, image, and semantics, forming Interest Units (IUs). Then we redesign the recommendation process into two stages. In the first stage, the focus is on recommend these Interest Units, capturing broad-level interests. In the second stage, it guides users to find the best option among similar products within the selected Interest Unit. User-IU interactions are incorporated into our ranking models, offering the advantage of more persistent IU behaviors compared to item-specific interactions. This interest unit based recommendation can be beneficial from mitigating the side effect of limited-stock problem, since most interaction can be gathered on interest units which can persist and accumulate over time. Experimental results on the production dataset and online A/B testing demonstrate the effectiveness and superiority of our proposed IU-centric recommendation approach. This study not only advances recommendation technologies but also emphasizes the potential for co-evolution between product innovations and the technologies involved in item supply and distribution. Jialiang Zhou, Qinye Xie, Qingheng Zhang, Yin Zhang 0006, Shuguang Han, Fei Huang 0002, Jufeng Chen |
KDD (2) | 8 |
| 2025 | RelationAdapter: Learning and Transferring Visual Relation with Diffusion TransformersabstractInspired by the in-context learning mechanism of large language models (LLMs), a new paradigm of generalizable visual prompt-based image editing is emerging. Existing single-reference methods typically focus on style or appearance adjustments and struggle with non-rigid transformations. To address these limitations, we propose leveraging source-target image pairs to extract and transfer content-aware editing intent to novel query images. To this end, we introduce RelationAdapter, a lightweight module that enables Diffusion Transformer (DiT) based models to effectively capture and apply visual transformations from minimal examples. We also introduce Relation252K, a comprehensive dataset comprising 218 diverse editing tasks, to evaluate model generalization and adaptability in visual prompt-driven scenarios. Experiments on Relation252K show that RelationAdapter significantly improves the model’s ability to understand and transfer editing intent, leading to notable gains in generation quality and overall editing performance. Yiren Song, Yicheng Li 0005, Yin Zhang 0006 |
NeurIPS | 5 |
| 2025 | Beyond decomposition: Hierarchical dependency management in multi-document question answeringabstractAbstract When using retrieval‐augmented generation (RAG) to handle multi‐document question answering (MDQA) tasks, it is beneficial to decompose complex queries into multiple simpler ones to enhance retrieval results. However, previous strategies always employ a one‐shot approach of question decomposition, overlooking subquestions dependency problem and failing to ensure that the derived subqueries are single‐hop. To overcome this challenge, we introduce a novel framework called DSRC‐QCS. Decompose‐solve‐renewal‐cycle (DSRC) is an iterative multi‐hop question processing module. The key idea of DSRC involves using a unique symbol to achieve hierarchical dependency management and employing a cyclical process of question decomposition, solving, and renewal to continuously generate and resolve all single‐hop subquestions. Query‐chain selector (QCS) functions as a voting mechanism that effectively utilizes the reasoning process of DSRC to assess and select solutions. We compare DSRC‐QCS against five RAG approaches across three datasets and three LLMs. DSRC‐QCS demonstrates superior performance. Compared to the Direct Retrieval method, DSRC‐QCS improves the average F1 score by 17.36% with Alpaca‐7b, 10.83% with LLaMa2‐Chat‐7b, and 11.88% with GPT‐3.5‐Turbo. We also conduct ablation studies to validate the performance of both DSRC and QCS and explore factors influencing the effectiveness of DSRC. We have included all prompts in the Appendix. Xiaoyan Zheng, Qianglong Chen, Yin Zhang 0006 |
J. Assoc. Inf. Sci. Technol. | 4 |
| 2025 | Critique of "Productivity, Portability, Performance: Data-Centric Python" by SCC Team From Zhejiang UniversityabstractIn SC'21, Alexandros Nikolaos Ziogas et al. proposed a Data-Centric Python workflow in their DaCe paper. DaCe provides high productivity, performance, and portability with language extensions and automatic optimizations. We reproduce the performance evaluation results from the paper on both CPU and GPU on the Azure CycleCloud cluster. We also reproduce the scaling results with up to 32 nodes and 64 processes. Our results show that the proposed workflow in that paper has outstanding performance and scalability in the provided cluster, in accordance with the SC paper. Zihan Yang 0004, Kaiqi Chen 0002, Xingjian Qian, Shaojun Xu, Chong Zeng 0001, Jianhai Chen, Yin Zhang 0006, Zeke Wang |
IEEE Trans. Parallel Distributed Syst. | 9 |
| 2024 | Towards Autonomous Tool Utilization in Language Models: A Unified, Efficient and Scalable FrameworkabstractIn recent research, significant advancements have been achieved in tool learning for large language models. Looking towards future advanced studies, the issue of fully autonomous tool utilization is particularly intriguing: given only a query, language models can autonomously decide whether to employ a tool, which specific tool to select, and how to utilize these tools, all without needing any tool-specific prompts within the context. To achieve this, we introduce a unified, efficient, and scalable framework for fine-tuning language models. Based on the degree of tool dependency, we initially categorize queries into three distinct types. By transforming the entire process into a sequential decision-making problem through conditional probability decomposition, our approach unifies the three types and autoregressively generates decision processes. Concurrently, we’ve introduced an “instruct, execute, and reformat” strategy specifically designed for efficient data annotation. Through end-to-end training on the annotated dataset comprising 26 diverse APIs, the model demonstrates a level of self-awareness, automatically seeking tool assistance when necessary. It significantly surpasses original instruction-tuned open-source language models and GPT-3.5/4 on multiple evaluation metrics. To address real-world scalability needs, we’ve enhanced our framework with a dynamic rehearsal strategy for continual learning, proven to require minimal new annotations to exhibit remarkable performance. Yicheng Li 0005, Hequan Ye, Yin Zhang 0006 |
LREC/COLING | 4 |
| 2024 | Teaching Small Language Models Reasoning through Counterfactual DistillationabstractWith the rise of large language models (LLMs), many studies are interested in transferring the reasoning capabilities of LLMs to small language models (SLMs).Previous distillation methods usually utilize the capabilities of LLMs to generate chain-of-thought (CoT) samples and teach SLMs via fine-tuning.However, such a standard distillation approach performs poorly when applied to out-of-distribution (OOD) examples, and the diversity of the generated CoT samples is insufficient.In this work, we propose a novel counterfactual distillation framework.Firstly, we leverage LLMs to automatically generate high-quality counterfactual data.Given an input text example, our method generates a counterfactual example that is very similar to the original input, but its task label has been changed to the desired one.Then, we utilize multi-view CoT to enhance the diversity of reasoning samples.Experiments on four NLP benchmarks show that our approach enhances the reasoning capabilities of SLMs and is more robust to OOD data.We also conduct extensive ablations and sample studies to understand the reasoning capabilities of SLMs. Yicheng Li 0005, Yin Zhang 0006 |
EMNLP | 6 |
| 2024 | DmRPC: Disaggregated Memory-aware Datacenter RPC for Data-intensive ApplicationsabstractModern datacenter applications are increasingly being built using a microservices architecture. These microservices communicate with each other using datacenter RPCs. RPC's pass by value semantics incur redundant data movement along the network, especially for data-intensive applications. Naively introducing a shared global address space to datacenter RPC does not work as it would couple microservices and require microservices to handle data consistency, significantly complicating the development and deployment of applications. Fortunately, the modern datacenter is embracing disaggregated memory (DM). In a DM-enabled datacenter, servers running the microservices can be all connected to one global disaggregated memory pool, thus the pass by value semantics can be replaced by pass by reference. However, prior work on DM requires complicated synchronization primitives to share data across physical machines, so naively adopting them to datacenter RPC would harm microservices' agility and modularity. To this end, we present DmRPC, a DM-aware datacenter RPC for data-intensive datacenter applications to our knowledge. First, DmRPC introduces a DM-aware shared global address space to provide the semantics of pass by reference to datacenter RPC, thus alleviating the redundant data movement issue. Second, DmRPC adopts a copy-on-write mechanism to avoid complicating application logic to handle data consistency while guaranteeing high performance. We have applied DmRPC to two different implementations of DM, one is network-based (DmRPC-net) while the other is CXL-based (DmRPC-CXL). Our evaluations on synthetic 7-tier microservices workloads show that DmRPC-net (or DmRPC-CXL) achieves 4.2× (or 8.3×) higher throughput and achieves 1.1 × (or 1.7 ×) lower average latency than that of the baseline, respectively. On a widely used microservice benchmark DeathStarBench, DmRPC-net can achieve 3.1 × higher throughput and 2.5 × lower average latency than the baseline. Jie Zhang 0081, Xuzheng Chen, Yin Zhang 0006, Zeke Wang |
ICDE | 3 |
| 2024 | Demystifying Datapath Accelerator Enhanced Off-path SmartNICabstractNetwork speeds grow quickly in the modern cloud, so SmartNICs are introduced to offload network processing tasks, even application logic. However, typical multicore SmartNICs such as BlueFiled-2 are only capable of processing control-plane tasks with their embedded processors that have limited memory bandwidth and computing power. On the other hand, cloud applications evolve rapidly, such that a limited number of fixed hardware engines in a SmartNIC cannot satisfy the requirements of cloud applications. Therefore, SmartNIC programmers call for a programmable datapath accelerator (DPA) to process network traffic at line rate. However, no existing work has unveiled the performance characteristics of the existing DPA. To this end, we present the first architectural characterization of the latest DPA-enhanced BlueFiled-3 (BF3) SmartNIC. Our evaluation results indicate that BF3's DPA is significantly wimpier than the off-path Arm processor and the host CPU. However, we still identify that DPA has three unique architectural characteristics that unleash the performance potential of DPA. Specifically, we demonstrate how to take advantage of DPA's three architectural characteristics regarding computing, networking, and memory subsystems. Then we propose three important guidelines for programmers to fully unleash the potential of DPA. To demonstrate the effectiveness of our approach, we conduct detailed case studies regarding each guideline. Our case study on key-value aggregation achieves up to$4.3 \times$higher throughput by using our guidelines to optimize memory combinations. Xuzheng Chen, Jie Zhang 0081, Lingjun Zhu, Yin Zhang 0006, Ming Liu 0027, Zeke Wang |
ICNP | 9 |
| 2024 | DMNet: Self-comparison Driven Model for Subject-independent Seizure DetectionabstractAutomated seizure detection (ASD) using intracranial electroencephalography (iEEG) is critical for effective epilepsy treatment. However, the significant domain shift of iEEG signals across subjects poses a major challenge, limiting their applicability in real-world clinical scenarios. In this paper, we address this issue by analyzing the primary cause behind the failure of existing iEEG models for subject-independent seizure detection, and identify a critical universal seizure pattern: seizure events consistently exhibit higher average amplitude compared to adjacent normal events. To mitigate the domain shifts and preserve the universal seizure patterns, we propose a novel self-comparison mechanism. This mechanism effectively aligns iEEG signals across subjects and time intervals. Building upon these findings, we propose Difference Matrix-based Neural Network (DMNet), a subject-independent seizure detection model, which leverages self-comparison based on two constructed (contextual, channel-level) references to mitigate shifts of iEEG, and utilize a simple yet effective difference matrix to encode the universal seizure patterns. Extensive experiments show that DMNet significantly outperforms previous SOTAs while maintaining high efficiency on a real-world clinical dataset collected by us and two public datasets for subject-independent seizure detection. Moreover, the visualization results demonstrate that the generated difference matrix can effectively capture the seizure activity changes during the seizure evolution process. Additionally, we deploy our method in an online diagnosis system to illustrate its effectiveness in real clinical applications. Shihao Tu, Linfeng Cao, Daoze Zhang, Lvbin Ma, Yin Zhang 0006, Yang Yang 0009 |
NeurIPS | 6 |
| 2024 | RLGAT: Retweet prediction in social networks using representation learning and GATs
Yin Zhang 0006, Shihua Cao |
Multim. Tools Appl. | 2 |
| 2023 | Task Difficulty Aware Parameter Allocation & Regularization for Lifelong LearningabstractParameter regularization or allocation methods are effective in overcoming catastrophic forgetting in lifelong learning. However, they solve all tasks in a sequence uniformly and ignore the differences in the learning difficulty of different tasks. So parameter regularization methods face significant forgetting when learning a new task very different from learned tasks, and parameter allocation methods face unnecessary parameter overhead when learning simple tasks. In this paper, we propose the Parameter Allocation & Regularization (PAR), which adaptively select an appropriate strategy for each task from parameter allocation and regularization based on its learning difficulty. A task is easy for a model that has learned tasks related to it and vice versa. We propose a divergence estimation method based on the Nearest-Prototype distance to measure the task relatedness using only features of the new task. Moreover, we propose a time-efficient relatedness-aware sampling-based architecture search strategy to reduce the parameter overhead for allocation. Experimental results on multiple benchmarks demonstrate that, compared with SOTAs, our method is scalable and significantly reduces the model's redundancy while improving the model's performance. Further qualitative analysis indicates that PAR obtains reasonable task-relatedness. Wenjin Wang 0003, Yunqing Hu, Qianglong Chen, Yin Zhang 0006 |
CVPR | 4 |
| 2023 | Dual Collaborative Visual-Semantic Mapping for Multi-Label Zero-Shot Image RecognitionabstractMulti-label zero-shot learning (ML-ZSL), with the difficulty of both multi-label learning and zero-shot learning, aims to recognize various unseen objects that are not observed during training. Previous methods mainly use a single directional visual-semantic mapping to associate the visual and semantic embedding space, which is not sufficient to adequately realize knowledge transfer from seen to unseen classes. In this paper, we propose a novel dual collaborative visual-semantic mapping framework, constructing abundant connection relationships by exploring two aspects of mapping streams, i.e., the visual-to-semantic (V2S) mapping and the semantic-to-visual (S2V) mapping. Through the collaborative learning of these two effective mappings, our method achieves state-of-the-art performance on the MS-COCO and PASCAL-VOC, two benchmarks for ML-ZSL. Yunqing Hu, Xuan Jin, Yin Zhang 0006 |
ICASSP | 4 |
| 2023 | A knowledge-guided and traditional Chinese medicine informed approach for herb recommendationabstractTraditional Chinese medicine (TCM) is an interesting research topic in China’s thousands of years of history. With the recent advances in artificial intelligence technology, some researchers have started to focus on learning the TCM prescriptions in a data-driven manner. This involves appropriately recommending a set of herbs based on patients’ symptoms. Most existing herb recommendation models disregard TCM domain knowledge, for example, the interactions between symptoms and herbs and the TCM-informed observations (i.e., TCM formulation of prescriptions). In this paper, we propose a knowledge-guided and TCM-informed approach for herb recommendation. The knowledge used includes path interactions and co-occurrence relationships among symptoms and herbs from a knowledge graph generated from TCM literature and prescriptions. The aforementioned knowledge is used to obtain the discriminative feature vectors of symptoms and herbs via a graph attention network. To increase the ability of herb prediction for the given symptoms, we introduce TCM-informed observations in the prediction layer. We apply our proposed model on a TCM prescription dataset, demonstrating significant improvements over state-of-the-art herb recommendation methods. Zhe Jin 0003, Yin Zhang 0006, Jiaxu Miao, Yi Yang 0001, Yueting Zhuang, Yunhe Pan |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2023 | Federated unsupervised representation learningabstractTo leverage the enormous amount of unlabeled data on distributed edge devices, we formulate a new problem in federated learning called federated unsupervised representation learning (FURL) to learn a common representation model without supervision while preserving data privacy. FURL poses two new challenges: (1) data distribution shift (non-independent and identically distributed, non-IID) among clients would make local models focus on different categories, leading to the inconsistency of representation spaces; (2) without unified information among the clients in FURL, the representations across clients would be misaligned. To address these challenges, we propose the federated contrastive averaging with dictionary and alignment (FedCA) algorithm. FedCA is composed of two key modules: a dictionary module to aggregate the representations of samples from each client which can be shared with all clients for consistency of representation space and an alignment module to align the representation of each client on a base model trained on public data. We adopt the contrastive approach for local model training. Through extensive experiments with three evaluation protocols in IID and non-IID settings, we demonstrate that FedCA outperforms all baselines with significant margins. Fengda Zhang, Kun Kuang 0001, Long Chen 0016, Zhaoyang You, Tao Shen 0002, Jun Xiao 0001, Yin Zhang 0006, Chao Wu 0001, Fei Wu 0001, Yueting Zhuang |
Frontiers Inf. Technol. Electron. Eng. | 7 |
| 2022 | Continual Few-shot Intent DetectionabstractIntent detection is at the core of task-oriented dialogue systems. Existing intent detection systems are typically trained with a large amount of data over a predefined set of intent classes. However, newly emerged intents in multiple domains are commonplace in the real world. And it is time-consuming and impractical for dialogue systems to re-collect enough annotated data and re-train the model. These limitations call for an intent detection system that could continually recognize new intents with very few labeled examples. In this work, we study the Continual Few-shot Intent Detection (CFID) problem and construct a benchmark consisting of nine tasks with multiple domains and imbalanced classes. To address the key challenges of (a) catastrophic forgetting during continuous learning and (b) negative knowledge transfer across tasks, we propose the Prefix-guided Lightweight Encoder (PLE) with three auxiliary strategies, namely Pseudo Samples Replay (PSR), Teacher Knowledge Transfer (TKT) and Dynamic Weighting Replay (DWR). Extensive experiments demonstrate the effectiveness and efficiency of our method in preventing catastrophic forgetting and encouraging positive knowledge transfer across tasks. Guodun Li, Yuchen Zhai, Qianglong Chen, Ji Zhang 0011, Yin Zhang 0006 |
COLING | 6 |
| 2022 | Diverse Instance Discovery: Vision-Transformer for Instance-Aware Multi-Label Image RecognitionabstractPrevious works on multi-label image recognition (MLIR) usually use CNNs as a starting point for research. In this paper, we take pure Vision Transformer (ViT) as the research base and make full use of the advantages of Transformer with long-range dependency modeling to circumvent the disadvantages of CNNs limited to local receptive field. However, for multi-label images containing multiple objects from different categories, scales, and spatial relations, it is not optimal to use global information alone. Our goal is to leverage ViT's patch tokens and self-attention mechanism to mine rich instances in multi-label images, named diverse instance discovery (DiD). To this end, we propose a semantic category-aware module and a spatial relationship-aware module, respectively, and then combine the two by a re-constraint strategy to obtain instance-aware attention maps. Finally, we propose a weakly supervised object localization-based approach to extract multi-scale local features, to form a multi-view pipeline. Our method requires only weakly supervised information at the label level, no additional knowledge injection or other strongly supervised information is required. Experiments on three benchmark datasets show that our method significantly outperforms previous works and achieves state-of-the-art results under fair experimental comparisons. Yunqing Hu, Xuan Jin, Yin Zhang 0006, Haiwen Hong, Jingfeng Zhang, Feihu Yan, Yuan He 0011, Hui Xue 0001 |
ICME | 3 |
| 2022 | DictBERT: Dictionary Description Knowledge Enhanced Language Model Pre-training via Contrastive LearningabstractAlthough pre-trained language models (PLMs) have achieved state-of-the-art performance on various natural language processing (NLP) tasks, they are shown to be lacking in knowledge when dealing with knowledge driven tasks. Despite the many efforts made for injecting knowledge into PLMs, this problem remains open. To address the challenge, we propose DictBERT, a novel approach that enhances PLMs with dictionary knowledge which is easier to acquire than knowledge graph (KG). During pre-training, we present two novel pre-training tasks to inject dictionary knowledge into PLMs via contrastive learning: dictionary entry prediction and entry description discrimination. In fine-tuning, we use the pre-trained DictBERT as a plugin knowledge base (KB) to retrieve implicit knowledge for identified entries in an input sequence, and infuse the retrieved knowledge into the input to enhance its representation via a novel extra-hop attention mechanism. We evaluate our approach on a variety of knowledge driven and language understanding tasks, including NER, relation extraction, CommonsenseQA, OpenBookQA and GLUE. Experimental results demonstrate that our model can significantly improve typical PLMs: it gains a substantial improvement of 0.5%, 2.9%, 9.0%, 7.1% and 3.3% on BERT-large respectively, and is also effective on RoBERTa-large. Qianglong Chen, Feng-Lin Li, Guohai Xu, Ming Yan 0008, Ji Zhang 0011, Yin Zhang 0006 |
IJCAI | 6 |
| 2022 | mmLayout: Multi-grained MultiModal Transformer for Document UnderstandingabstractRecent efforts of multimodal Transformers have improved Visually Rich Document Understanding (VrDU) tasks via incorporating visual and textual information. However, existing approaches mainly focus on fine-grained elements such as words and document image patches, making it hard for them to learn from coarse-grained elements, including natural lexical units like phrases and salient visual regions like prominent image regions. In this paper, we attach more importance to coarse-grained elements containing high-density information and consistent semantics, which are valuable for document understanding. At first, a document graph is proposed to model complex relationships among multi-grained multimodal elements, in which salient visual regions are detected by a cluster-based method. Then, a multi-grained multimodal Transformer called mmLayout is proposed to incorporate coarse-grained information into existing pre-trained fine-grained multimodal Transformers based on the graph. In mmLayout, coarse-grained information is aggregated from fine-grained, and then, after further processing, is fused back into fine-grained for final prediction. Furthermore, common sense enhancement is introduced to exploit the semantic information of natural lexical units. Experimental results on four tasks, including information extraction and document question answering, show that our method can improve the performance of multimodal Transformers based on fine-grained elements and achieve better performance with fewer parameters. Qualitative analyses show that our method can capture consistent semantics in coarse-grained elements. Wenjin Wang 0003, Zhengjie Huang, Qianglong Chen, Qiming Peng, Yinxu Pan, Weichong Yin, Shikun Feng, Yu Sun 0029, Dianhai Yu, Yin Zhang 0006 |
ACM Multimedia | 11 |
| 2022 | Align and Adapt: A Two-stage Adaptation Framework for Unsupervised Domain AdaptationabstractUnsupervised domain adaptation aims to transfer knowledge from a labeled but heterogeneous source domain to an unlabeled target domain, alleviating the labeling efforts. Early advances in domain adaptation focus on invariant representations learning (IRL) methods to align domain distributions. Recent studies further utilize semi-supervised learning (SSL) methods to regularize domain-invariant representations based on the cluster assumption, making the category boundary more clear. However, the misalignment in the IRL methods might be intensified by SSL methods if the target instances are more proximate to the wrong source centroid, resulting in incompatibility between these techniques. In this paper, we hypothesize this phenomenon derives from the distraction of the source domain, and further give a novel two-stage adaptation framework to adapt the model toward the target domain. In addition, we propose DCAN to reduce the misalignment in IRL methods in the first stage, and we propose PCST to encode the semantic structure of unlabeled target data in the second stage. Extensive experiments demonstrate that our method outperforms current state-of-the-art methods on four benchmarks (Office-31, ImageCLEF-DA, Office-Home, and VisDA-2017). Yuchen Zhai, Yin Zhang 0006 |
ACM Multimedia | 3 |
| 2022 | Rethinking the Value of Gazetteer in Chinese Named Entity Recognition
Qianglong Chen, Xiangji Zeng, Jiangang Zhu, Yin Zhang 0006, Bojia Lin, Yang Yang 0012, Daxin Jiang |
NLPCC (1) | 4 |
| 2022 | DAMO-NLP at NLPCC-2022 Task 2: Knowledge Enhanced Robust NER for Speech Entity Linking
Shen Huang, Yuchen Zhai, Xinwei Long, Yong Jiang 0005, Xiaobin Wang, Yin Zhang 0006, Pengjun Xie |
NLPCC (2) | 6 |
| 2022 | Multi-label Masked Language Modeling on Zero-shot Code-switched Sentiment AnalysisabstractIn multilingual communities, code-switching is a common phenomenon and code-switched tasks have become a crucial area of research in natural language processing (NLP) applications. Existing approaches mainly focus on supervised learning. However, it is expensive to annotate a sufficient amount of code-switched data. In this paper, we consider zero-shot setting and improve model performance on code-switched tasks via monolingual language datasets, unlabeled code-switched datasets, and semantic dictionaries. Inspired by the mechanism of code-switching itself, we propose multi-label masked language modeling and predict both the masked word and its synonyms in other languages. Experimental results show that compared with baselines, our method can further improve the pretrained multilingual model's performance on code-switched sentiment analysis datasets. Ji Zhang 0011, Yin Zhang 0006 |
SIGIR | 4 |
| 2022 | FEUI: Fusion Embedding for User Identification across social networks
Yin Zhang 0006, Keyong Hu |
Appl. Intell. | 2 |
| 2022 | NGAT: attention in breadth and depth exploration for semi-supervised graph representation learningabstractRecently, graph neural networks (GNNs) have achieved remarkable performance in representation learning on graph-structured data. However, as the number of network layers increases, GNNs based on the neighborhood aggregation strategy deteriorate due to the problem of oversmoothing, which is the major bottleneck for applying GNNs to real-world graphs. Many efforts have been made to improve the process of feature information aggregation from directly connected nodes, i.e., breadth exploration. However, these models perform the best only in the case of three or fewer layers, and the performance drops rapidly for deep layers. To alleviate oversmoothing, we propose a nested graph attention network (NGAT), which can work in a semi-supervised manner. In addition to breadth exploration, a k -layer NGAT uses a layer-wise aggregation strategy guided by the attention mechanism to selectively leverage feature information from the k th -order neighborhood, i.e., depth exploration. Even with a 10-layer or deeper architecture, NGAT can balance the need for preserving the locality (including root node features and the local structure) and aggregating the information from a large neighborhood. In a number of experiments on standard node classification tasks, NGAT outperforms other novel models and achieves state-of-the-art performance. Jianke Hu, Yin Zhang 0006 |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2022 | FEBDNN: fusion embedding-based deep neural network for user retweeting behavior prediction on social networks
Yin Zhang 0006, Keyong Hu, Shihua Cao |
Neural Comput. Appl. | 2 |
| 2022 | XCode: Towards Cross-Language Code Representation with Large-Scale Pre-TrainingabstractSource code representation learning is the basis of applying artificial intelligence to many software engineering tasks such as code clone detection, algorithm classification, and code summarization. Recently, many works have tried to improve the performance of source code representation from various perspectives, e.g., introducing the structural information of programs into latent representation. However, when dealing with rapidly expanded unlabeled cross-language source code datasets from the Internet, there are still two issues. Firstly, deep learning models for many code-specific tasks still suffer from the lack of high-quality labels. Secondly, the structural differences among programming languages make it more difficult to process multiple languages in a single neural architecture. To address these issues, in this article, we propose a novel Cross -language Code representation with a large-scale pre-training ( XCode ) method. Concretely, we propose to use several abstract syntax trees and ELMo-enhanced variational autoencoders to obtain multiple pre-trained source code language models trained on about 1.5 million code snippets. To fully utilize the knowledge across programming languages, we further propose a Shared Encoder-Decoder (SED) architecture which uses the multi-teacher single-student method to transfer knowledge from the aforementioned pre-trained models to the distilled SED. The pre-trained models and SED will cooperate to better represent the source code. For evaluation, we examine our approach on three typical downstream cross-language tasks, i.e., source code translation, code clone detection, and code-to-code search, on a real-world dataset composed of programming exercises with multiple solutions. Experimental results demonstrate the effectiveness of our proposed approach on cross-language code representations. Meanwhile, our approach performs significantly better than several code representation baselines on different downstream tasks in terms of multiple automatic evaluation metrics. Zehao Lin, Guodun Li, Jingfeng Zhang, Xiangji Zeng, Yin Zhang 0006, Yao Wan 0001 |
ACM Trans. Softw. Eng. Methodol. | 6 |
| 2021 | Reinforced History Backtracking for Conversational Question AnsweringabstractTo model the context history in multi-turn conversations has become a critical step towards a better understanding of the user query in question answering systems. To utilize the context history, most existing studies treat the whole context as input, which will inevitably face the following two challenges. First, modeling a long history can be costly as it requires more computation resources. Second, the long context history consists of a lot of irrelevant information that makes it difficult to model appropriate information relevant to the user query. To alleviate these problems, we propose a reinforcement learning based method to capture and backtrack the related conversation history to boost model performance in this paper. Our method seeks to automatically backtrack the history information with the implicit feedback from the model performance. We further consider both immediate and delayed rewards to guide the reinforced backtracking policy. Extensive experiments on a large conversational question answering dataset show that the proposed method can help to alleviate the problems arising from longer context history. Meanwhile, experiments show that the method yields better performance than other strong baselines, and the actions made by the method are insightful. Minghui Qiu, Xinjing Huang, Cen Chen 0001, Chen Qu 0001, Wei Wei 0002, Jun Huang 0007, Yin Zhang 0006 |
AAAI | 8 |
| 2021 | KACE: Generating Knowledge Aware Contrastive Explanations for Natural Language InferenceabstractQianglong Chen, Feng Ji, Xiangji Zeng, Feng-Lin Li, Ji Zhang, Haiqing Chen, Yin Zhang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Qianglong Chen, Xiangji Zeng, Feng-Lin Li, Ji Zhang 0011, Haiqing Chen, Yin Zhang 0006 |
ACL/IJCNLP (1) | 7 |
| 2021 | CAT-BERT: A Context-Aware Transferable BERT Model for Multi-turn Machine Reading Comprehension
Cen Chen 0001, Xinjing Huang, Chengyu Wang 0001, Minghui Qiu, Jun Huang 0007, Yin Zhang 0006 |
DASFAA (2) | 7 |
| 2021 | Meta Distant Transfer Learning for Pre-trained Language ModelsabstractWith the wide availability of Pre-trained Language Models (PLMs), multi-task fine-tuning across domains has been extensively applied.For tasks related to distant domains with different class label sets, PLMs may memorize nontransferable knowledge for the target domain and suffer from negative transfer.Inspired by meta-learning, we propose the Meta Distant Transfer Learning (Meta-DTL) framework to learn the cross-task knowledge for PLM-based methods.Meta-DTL first employs task representation learning to mine implicit relations among multiple tasks and classes.Based on the results, it trains a PLM-based meta-learner to capture the transferable knowledge across tasks.The weighted maximum entropy regularizers are proposed to make meta-learner more task-agnostic and unbiased.Finally, the meta-learner can be fine-tuned to fit each task with better parameter initialization.We evaluate Meta-DTL using both BERT and ALBERT on seven public datasets.Experiment results confirm the superiority of Meta-DTL as it consistently outperforms strong baselines.We find that Meta-DTL is highly effective when very few data is available for the target task. Chengyu Wang 0001, Haojie Pan, Minghui Qiu, Jun Huang 0007, Fei Yang 0007, Yin Zhang 0006 |
EMNLP (1) | 6 |
| 2021 | Fix-Filter-Fix: Intuitively Connect Any Models for Effective Bug FixingabstractLocating and fixing bugs is a time-consuming task.Most neural machine translation (NMT) based approaches for automatically bug fixing lack generality and do not make full use of the rich information in the source code.In NMTbased bug fixing, we find some predicted code identical to the input buggy code (called unchanged fix) in NMT-based approaches due to high similarity between buggy and fixed code (e.g., the difference may only appear in one particular line).Obviously, unchanged fix is not the correct fix because it is the same as the buggy code that needs to be fixed.Based on these, we propose an intuitive yet effective general framework (called Fix-Filter-Fix or F 3 ) for bug fixing.F 3 connects models with our filter mechanism to filter out the last model's unchanged fix to the next.We propose an F 3 theory that can quantitatively and accurately calculate the F 3 lifting effect.To evaluate, we implement the Seq2Seq Transformer (ST) and the AST2Seq Transformer (AT) to form some basic F 3 instances, called F 3 ST +AT and F 3 AT +ST .Comparing them with single model approaches and many model connection baselines across four datasets validates the effectiveness and generality of F 3 and corroborates our findings and methodology. Haiwen Hong, Jingfeng Zhang, Yin Zhang 0006, Yao Wan 0001, Yulei Sui |
EMNLP (1) | 3 |
| 2021 | DRDF: Determining the Importance of Different Multimodal Information with Dual-Router Dynamic FrameworkabstractIn multimodal tasks, the importance of text and image modal information often varies for different input cases. To model the difference of importance of different modal information, we propose a high-performance and highly general Dual-Router Dynamic Framework (DRDF), consisting of Dual-Router, MWF-Layer, experts and expert fusion unit. The text router and image router in Dual-Router take text modal information and image modal information respectively, and MWF-Layer is responsible to determine the importance of modal information. Based on the result of the determination, MWF-Layer generates fused weights for the subsequent experts fusion. Experts can adopt a variety of backbones that match the current multimodal or unimodal task. DRDF features high generality and modularity, and we test 12 backbones such as Visual BERT and their corresponding DRDF instances on the multimodal dataset Hateful memes, and unimodal datasets CIFAR10, CIFAR100, and TinyImagenet. Our DRDF instance outperforms those backbones. We also validate the effectiveness of components of DRDF by ablation studies, and discuss the reasons and ideas of DRDF design. Haiwen Hong, Xuan Jin, Yin Zhang 0006, Yunqing Hu, Jingfeng Zhang, Yuan He 0011, Hui Xue 0001 |
ACM Multimedia | 3 |
| 2021 | RAMS-Trans: Recurrent Attention Multi-scale Transformer for Fine-grained Image RecognitionabstractIn fine-grained image recognition (FGIR), the localization and amplification of region attention is an important factor, which has been explored extensively convolutional neural networks (CNNs) based approaches. The recently developed vision transformer (ViT) has achieved promising results in computer vision tasks. Compared with CNNs, Image sequentialization is a brand new manner. However, ViT is limited in its receptive field size and thus lacks local attention like CNNs due to the fixed size of its patches, and is unable to generate multi-scale features to learn discriminative region attention. To facilitate the learning of discriminative region attention without box/part annotations, we use the strength of the attention weights to measure the importance of the patch tokens corresponding to the raw images. We propose the recurrent attention multi-scale transformer (RAMS-Trans), which uses the transformer's self-attention to recursively learn discriminative region attention in a multi-scale manner. Specifically, at the core of our approach lies the dynamic patch proposal module (DPPM) responsible for guiding region amplification to complete the integration of multi-scale image patches. The DPPM starts with the full-size image patches and iteratively scales up the region attention to generate new patches from global to local by the intensity of the attention weights generated at each scale as an indicator. Our approach requires only the attention weights that come with ViT itself and can be easily trained end-to-end. Extensive experiments demonstrate that RAMS-Trans performs better than exising works, in addition to efficient CNN models, achieving state-of-the-art results on three benchmark datasets. Yunqing Hu, Xuan Jin, Yin Zhang 0006, Haiwen Hong, Jingfeng Zhang, Yuan He 0011, Hui Xue 0001 |
ACM Multimedia | 3 |
| 2021 | Similar Scenes Arouse Similar Emotions: Parallel Data Augmentation for Stylized Image CaptioningabstractStylized image captioning systems aim to generate a caption not only semantically related to a given image but also consistent with a given style description. One of the biggest challenges with this task is the lack of sufficient paired stylized data. Many studies focus on unsupervised approaches, without considering from the perspective of data augmentation. We begin with the observation that people may recall similar emotions when they are in similar scenes, and often express similar emotions with similar style phrases, which underpins our data augmentation idea. In this paper, we propose a novel Extract-Retrieve-Generate data augmentation framework to extract style phrases from small-scale stylized sentences and graft them to large-scale factual captions. First, we design the emotional signal extractor to extract style phrases from small-scale stylized sentences. Second, we construct the plugable multi-modal scene retriever to retrieve scenes represented with pairs of an image and its stylized caption, which are similar to the query image or caption in the large-scale factual data. In the end, based on the style phrases of similar scenes and the factual description of the current scene, we build the emotion-aware caption generator to generate fluent and diversified stylized captions for the current scene. Extensive experimental results show that our framework can alleviate the data scarcity problem effectively. It also significantly boosts the performance of several existing image captioning models in both supervised and unsupervised settings, which outperforms the state-of-the-art stylized image captioning methods in terms of both sentence relevance and stylishness by a substantial margin. Guodun Li, Yuchen Zhai, Zehao Lin, Yin Zhang 0006 |
ACM Multimedia | 4 |
| 2021 | Dialogue State Tracking with Multi-Level Fusion of Predicted Dialogue States and ConversationsabstractMost recently proposed approaches in dialogue state tracking (DST) leverage the context and the last dialogue states to track current dialogue states, which are often slot-value pairs.Although the context contains the complete dialogue information, the information is usually indirect and even requires reasoning to obtain.The information in the lastly predicted dialogue states is direct, but when there is a prediction error, the dialogue information from this source will be incomplete or erroneous.In this paper, we propose the Dialogue State Tracking with Multi-Level Fusion of Predicted Dialogue States and Conversations network (FPDSC).This model extracts information of each dialogue turn by modeling interactions among each turn utterance, the corresponding last dialogue states, and dialogue slots.Then the representation of each dialogue turn is aggregated by a hierarchical structure to form the passage information, which is utilized in the current turn of DST.Experimental results validate the effectiveness of the fusion network with 55.03% and 59.07%joint accuracy on MultiWOZ 2.0 and MultiWOZ 2.1 datasets, which reaches the state-of-the-art performance.Furthermore, we conduct the deleted-value and related-slot experiments on MultiWOZ 2.1 to evaluate our model. Jingyao Zhou, Haipang Wu, Zehao Lin, Guodun Li, Yin Zhang 0006 |
SIGDIAL | 5 |
| 2021 | Predict-Then-Decide: A Predictive Approach for Wait or Answer Task in Dialogue SystemsabstractDifferent people have different habits of describing their intents in conversations. Some people tend to deliberate their intents in several successive utterances, i.e., they use several consistent messages for readability instead of a long sentence to express their question. This creates a predicament faced by the application of dialogue systems, especially in real-world industry scenarios, in which the dialogue system is unsure whether it should answer the user's query immediately or wait for user's further supplementary input. Inspired by such interesting quandary, we define a novel task: Wait-or-Answer to better tackle this dilemma faced by dialogue systems. We shed light on a new research topic about how the dialogue system can be more intelligent to behave in this Wait-or-Answer quandary. Further, we propose a predictive approach dubbed Predict-then-Decide (PTD) to resolve this Wait-or-Answer task. More specifically, we take advantage of a decision model to help the dialogue system decide to wait or answer. The decision of decision model is made with the assistance of two ancillary prediction models: a user prediction and an agent prediction. The user prediction model tries to predict what the user would supplement and uses its prediction to persuade the decision model that the user has some information to add, so the dialogue system should wait. The agent prediction model, nevertheless, struggles to predict the answer of the dialogue system and convince the decision model that it's a superior choice to answer the user's query immediately since the user's input is finished. We conduct our experiment on two real-life scenarios and three public datasets.Experimental results on five datasets show our proposed PTD approach significantly outperforms the existing models in solving this Wait-or-Answer problem. Zehao Lin, Shaobo Cui 0001, Guodun Li, Xiaoming Kang, Feng-Lin Li, Zhongzhou Zhao, Haiqing Chen, Yin Zhang 0006 |
IEEE ACM Trans. Audio Speech Lang. Process. | 9 |
| 2021 | Local-Global Memory Neural Network for Medication PredictionabstractElectronic medical records (EMRs) play an important role in medical data mining and sequential data learning. In this article, we propose to use a sequential neural network with dynamic content-based memories to predict future medications, given EMRs. The local-global memory neural network contains two layers of memories: the local memory and the global memory. Particularly, our method learns the hidden knowledge within EMRs by locally remembering individual patterns of a patient (via local memory) and globally remembering group evidence of disease (via global memory). In addition, we show how our model can be modified to classify the hidden states of EMRs from different patients at each time step into different phases that indicate the progressions of medications in terms of a specific disease, in an unsupervised manner. Experimental results on real EMRs data sets show that, by learning EMRs with external local and global memories, with regard to a given disease, our model improves the prediction performance compared with several alternative methods. Jun Song 0004, Siliang Tang, Yin Zhang 0006, Zhigang Chen 0003, Zhongfei Zhang, Tong Zhang 0001, Fei Wu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2020 | MTSS: Learn from Multiple Domain Teachers and Become a Multi-Domain Dialogue ExpertabstractHow to build a high-quality multi-domain dialogue system is a challenging work due to its complicated and entangled dialogue state space among each domain, which seriously limits the quality of dialogue policy, and further affects the generated response. In this paper, we propose a novel method to acquire a satisfying policy and subtly circumvent the knotty dialogue state representation problem in the multi-domain setting. Inspired by real school teaching scenarios, our method is composed of multiple domain-specific teachers and a universal student. Each individual teacher only focuses on one specific domain and learns its corresponding domain knowledge and dialogue policy based on a precisely extracted single domain dialogue state representation. Then, these domain-specific teachers impart their domain knowledge and policies to a universal student model and collectively make this student model a multi-domain dialogue expert. Experiment results show that our method reaches competitive results with SOTAs in both multi-domain and single domain setting. Shuke Peng, Zehao Lin, Shaobo Cui 0001, Haiqing Chen, Yin Zhang 0006 |
AAAI | 6 |
| 2020 | Improving Commonsense Question Answering by Graph-based Iterative Retrieval over Multiple Knowledge SourcesabstractIn order to facilitate natural language understanding, the key is to engage commonsense or background knowledge.However, how to engage commonsense effectively in question answering systems is still under exploration in both research academia and industry.In this paper, we propose a novel question-answering method by integrating multiple knowledge sources, i.e.Con-ceptNet, Wikipedia, and the Cambridge Dictionary, to boost the performance.More concretely, we first introduce a novel graph-based iterative knowledge retrieval module, which iteratively retrieves concepts and entities related to the given question and its choices from multiple knowledge sources.Afterward, we use a pre-trained language model to encode the question, retrieved knowledge and choices, and propose an answer choice-aware attention mechanism to fuse all hidden representations of the previous modules.Finally, the linear classifier for specific tasks is used to predict the answer.Experimental results on the CommonsenseQA dataset show that our method significantly outperforms other competitive methods and achieves the new state-ofthe-art.In addition, further ablation studies demonstrate the effectiveness of our graph-based iterative knowledge retrieval module and the answer choice-aware attention module in retrieving and synthesizing background knowledge from multiple knowledge sources. Qianglong Chen, Haiqing Chen, Yin Zhang 0006 |
COLING | 4 |
| 2020 | Counterfactual Generator: A Weakly-Supervised Method for Named Entity RecognitionabstractPast progress on neural models has proven that named entity recognition is no longer a problem if we have enough labeled data.However, collecting enough data and annotating them are labor-intensive, time-consuming, and expensive.In this paper, we decompose the sentence into two parts: entity and context, and rethink the relationship between them and model performance from a causal perspective.Based on this, we propose the Counterfactual Generator, which generates counterfactual examples by the interventions on the existing observational examples to enhance the original dataset.Experiments across three datasets show that our method improves the generalization ability of models under limited observational examples.Besides, we provide a theoretical foundation by using a structural causal model to explore the spurious correlations between input features and output labels.We investigate the causal effects of entity or context on model performance under both conditions: the non-augmented and the augmented.Interestingly, we find that the non-spurious correlations are more located in entity representation rather than context representation.As a result, our method eliminates part of the spurious correlations between context representation and output labels.The code is available at https://github.com/xijiz/cfgen. Xiangji Zeng, Yunliang Li, Yuchen Zhai, Yin Zhang 0006 |
EMNLP (1) | 4 |
| 2020 | Hybrid embedding and joint training of stacked encoder for opinion question machine reading comprehensionabstractOpinion question machine reading comprehension (MRC) requires a machine to answer questions by analyzing corresponding passages. Compared with traditional MRC tasks where the answer to every question is a segment of text in corresponding passages, opinion question MRC is more challenging because the answer to an opinion question may not appear in corresponding passages but needs to be deduced from multiple sentences. In this study, a novel framework based on neural networks is proposed to address such problems, in which a new hybrid embedding training method combining text features is used. Furthermore, extra attention and output layers which generate auxiliary losses are introduced to jointly train the stacked recurrent neural networks. To deal with imbalance of the dataset, irrelevancy of question and passage is used for data augmentation. Experimental results show that the proposed method achieves state-of-the-art performance. We are the biweekly champion in the opinion question MRC task in Artificial Intelligence Challenger 2018 (AIC2018). Xiangzhou Huang, Siliang Tang, Yin Zhang 0006, Baogang Wei |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2019 | Task-Oriented Conversation Generation Using Heterogeneous Memory NetworksabstractZehao Lin, Xinjing Huang, Feng Ji, Haiqing Chen, Yin Zhang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Zehao Lin, Xinjing Huang, Haiqing Chen, Yin Zhang 0006 |
EMNLP/IJCNLP (1) | 5 |
| 2019 | Named Entity Recognition in Traditional Chinese Medicine Clinical Cases Combining BiLSTM-CRF with Knowledge Graph
Zhe Jin 0003, Yin Zhang 0006, Haodan Kuang, Yunhe Pan |
KSEM (1) | 2 |
| 2019 | Traditional Chinese medicine clinical records classification with BERT and domain specific corporaabstractTraditional Chinese Medicine (TCM) has been developed for several thousand years and plays a significant role in health care for Chinese people. This paper studies the problem of classifying TCM clinical records into 5 main disease categories in TCM. We explored a number of state-of-the-art deep learning models and found that the recent Bidirectional Encoder Representations from Transformers can achieve better results than other deep learning models and other state-of-the-art methods. We further utilized an unlabeled clinical corpus to fine-tune the BERT language model before training the text classifier. The method only uses Chinese characters in clinical text as input without preprocessing or feature engineering. We evaluated deep learning models and traditional text classifiers on a benchmark data set. Our method achieves a state-of-the-art accuracy 89.39% ± 0.35%, Macro F1 score 88.64% ± 0.40% and Micro F1 score 89.39% ± 0.35%. We also visualized attention weights in our method, which can reveal indicative characters in clinical text. Zhe Jin 0003, Chengsheng Mao, Yin Zhang 0006, Yuan Luo 0001 |
J. Am. Medical Informatics Assoc. | 4 |
| 2018 | Temporality-enhanced knowledgememory network for factoid question answeringabstractQuestion answering is an important problem that aims to deliver specific answers to questions posed by humans in natural language. How to efficiently identify the exact answer with respect to a given question has become an active line of research. Previous approaches in factoid question answering tasks typically focus on modeling the semantic relevance or syntactic relationship between a given question and its corresponding answer. Most of these models suffer when a question contains very little content that is indicative of the answer. In this paper, we devise an architecture named the temporality-enhanced knowledge memory network (TE-KMN) and apply the model to a factoid question answering dataset from a trivia competition called quiz bowl. Unlike most of the existing approaches, our model encodes not only the content of questions and answers, but also the temporal cues in a sequence of ordered sentences which gradually remark the answer. Moreover, our model collaboratively uses external knowledge for a better understanding of a given question. The experimental results demonstrate that our method achieves better performance than several state-of-the-art methods. Xinyu Duan, Siliang Tang, Shengyu Zhang 0001, Yin Zhang 0006, Zhou Zhao 0001, Jianru Xue, Yueting Zhuang, Fei Wu 0001 |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2018 | Image-based 3D model retrieval using manifold learningabstractWe propose a new framework for image-based three-dimensional (3D) model retrieval. We first model the query image as a Euclidean point. Then we model all projected views of a 3D model as a symmetric positive definite (SPD) matrix, which is a point on a Riemannian manifold. Thus, the image-based 3D model retrieval is reduced to a problem of Euclid-to-Riemann metric learning. To solve this heterogeneous matching problem, we map the Euclidean space and SPD Riemannian manifold to the same high-dimensional Hilbert space, thus shrinking the great gap between them. Finally, we design an optimization algorithm to learn a metric in this Hilbert space using a kernel trick. Any new image descriptors, such as the features from deep learning, can be easily embedded in our framework. Experimental results show the advantages of our approach over the state-of-the-art methods for image-based 3D model retrieval. Pan-pan Mu, Sanyuan Zhang, Yin Zhang 0006, Xiuzi Ye |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2018 | A Topic Modeling Approach for Traditional Chinese Medicine PrescriptionsabstractIn traditional Chinese medicine (TCM), prescriptions are the daughters of doctors' clinical experiences, which have been the main way to cure diseases in China for several thousand years. In the long Chinese history, a large number of prescriptions have been invented based on TCM theories. Regularities in the prescriptions are important for both clinical practice and novel prescription development. Previous works used many methods to discover regularities in prescriptions, but rarely described how a prescription is generated using TCM theories. In this work, we propose a topic model which characterizes the generative process of prescriptions in TCM theories and further incorporate domain knowledge into the topic model. Using 33,765 prescriptions in TCM prescription books, the model can reflect the prescribing patterns in TCM. Our method can outperform several previous topic models and group recommendation methods on generalization performance, herbs recommendation, symptoms suggestion, and prescribing patterns discovery. Yin Zhang 0006, Baogang Wei, Zhe Jin 0003 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2017 | Incorporating Knowledge Graph Embeddings into Topic ModelingabstractProbabilistic topic models could be used to extract low-dimension topics from document collections. However, such models without any human knowledge often produce topics that are not interpretable. In recent years, a number of knowledge-based topic models have been proposed, but they could not process fact-oriented triple knowledge in knowledge graphs. Knowledge graph embeddings, on the other hand, automatically capture relations between entities in knowledge graphs. In this paper, we propose a novel knowledge-based topic model by incorporating knowledge graph embeddings into topic modeling. By combining latent Dirichlet allocation, a widely used topic model with knowledge encoded by entity vectors, we improve the semantic coherence significantly and capture a better representation of a document in the topic space. Our evaluation results will demonstrate the effectiveness of our method. Yin Zhang 0006, Baogang Wei, Zhe Jin 0003, Qinfei Chen |
AAAI | 2 |
| 2017 | Detecting Temporal Proposal for Action Localization with Tree-structured Search PolicyabstractUnderstanding the semantics in videos is a complex but crucial task in video analysis. This paper focuses on localizing category-independent events, actions or other semantics in an untrimmed video, referred as salient temporal proposal localization. Traditional methods like sliding window have a high computational cost due to the densely sampling of different video segments. We propose a reinforcement learning based method, which trains a localizer that learns a search policy that, instead of exploring every video segment, finds an optimal search path to locate a salient proposal based on the currently observing video segment in a tree structure, therefore reduces the number of video segments fed into the proposal detector. In each search step, a localizer is trained to iteratively select the next sub-region containing salient proposals to continue the search, and a proposal detector is trained to recognize salient proposal from the sub-regions. The experiments demonstrate that our method is able to precisely detect salient proposals with a comparable recall and with much fewer candidate windows. Xinyang Jiang, Siliang Tang, Yang Yang 0009, Zhou Zhao 0001, Yin Zhang 0006, Fei Wu 0001, Yueting Zhuang |
ACM Multimedia | 5 |
| 2017 | Mining coherent topics in documents using word embeddings and large-scale text data
Yin Zhang 0006, Qinfei Chen, Hongze Qian, Baogang Wei, Zhifeng Hu |
Eng. Appl. Artif. Intell. | 2 |
| 2017 | Temporal Interaction and Causal Influence in Community-Based Question AnsweringabstractDuring the last decade, community-based question answering (CQA) sites have accumulated a vast amount of questions and their crowdsourced answers over time. How to efficiently identify the quality of answers that are relevant to a given question has become an active line of research in CQA. The major challenge of CQA is the accurate selection of high-quality answers w.r.t given questions. Previous approaches tend to model the semantic matching between individual pair of one question and its corresponding answer (how fitting an answer is to a posted question). However, these works ignore the temporal interactions between answers (how previous answers influence the late posted answers). For example, a rational user likely adapts others' opinions, revises his inclinations, and posts a more appropriate answer after understanding the given question and previously posted answers. As a result, this paper devises an architecture named Temporal Interaction and Causal Influence LSTM (TC-LSTM) to effectively leverage not only the causal influence between question-answer (how appropriate an answer is for a given question) but also the temporal interactions between answers-answer (how a high-quality answer gradually forms). In particular, long short-term memory (LSTM) is used to capture the explicit question-answer influence and the implicit answers-answer interactions. Experiments are conducted on SemEval 2015 CQA dataset for answer classification task and Baidu Zhidao Dataset for answer ranking task. The experimental results show the advantage of our model comparing with other state-of-the-art methods. Fei Wu 0001, Xinyu Duan, Jun Xiao 0001, Zhou Zhao 0001, Siliang Tang, Yin Zhang 0006, Yueting Zhuang |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2016 | A joint model for question-answering over Traditional Chinese MedicineabstractTraditional Chinese Medicine (TCM) has been around for over 2000 years and it's a significant part of Chinese cultural heritage. The theoretical framework of TCM is unique and with rich of content, which contains the complex relationships between disease and medicine and has formed a unique system to diagnose and cure illness. Research on question-answering (QA) over TCM is significant for Chinese NLP and representative, because the resources of TCM are mostly Chinese-based. The general strategies for QA are pipeline, includes Natural Language Processing (NLP) for original questions and structural query process. However, noises are easily introduced and information contained in the original questions may loss during these two separate procedures. In this paper we present a novel joint model for QA over TCM, which unites these two procedures into a uniform framework, together with Wikipedia alignment. The results of our evaluation with 100 benchmark queries demonstrate the value of our approach. Xiangzhou Huang, Yin Zhang 0006, Baogang Wei |
BIBM | 2 |
| 2016 | Traditional Chinese medicine clinical records classification using knowledge-powered document embeddingabstractText classification is one of the fundamental tasks in text mining. In the medical domain, there have been a number of studies on text classification in modern medicine clinical notes written in English. However, very limited text classification research has been conducted on clinical notes written in Chinese, especially traditional Chinese medicine (TCM) clinical records. The goal of this study was to investigate features and machine learning classification algorithms for TCM clinical text classification. We collected 7,037 TCM clinical records of famous TCM doctors as our dataset, and investigated the effects of different types of features and classification algorithms. Additionally, we proposed a novel method to combine deep learning text representation with TCM domain knowledge, which results in the best classification performance. Yin Zhang 0006, Baogang Wei, Zherong Li, Xiangzhou Huang |
BIBM | 2 |
| 2016 | Purchase prediction using Tmall-specific featuresabstractSummary Historical user activity, such as online shopping recommendations, content personalization, and advertising clicking rates, is a critical component of building user profiles to predict purchases and preferences. Alongside the rapid development of e‐commerce, purchase prediction has become an increasingly important consideration for a wide variety of retail platforms. This paper proposes a framework which combines machine learning methods with a threshold‐moving approach to predict sets of pairs (user id and brand id,) in terms of whether a certain brand is purchased by a specified user according to his or her historical activity records. Three specific feature groups are extracted: click features, purchase features, and collect‐and‐cart features using a dataset from Tmall, a Chinese business‐to‐consumer online retail platform. Next, seven user purchase prediction experiments with different combinations of the three feature groups are conducted, and the purchase prediction performance is observed. Results showed that a combination of all three feature groups, with 27 features in total, provided valuable purchase prediction contributions and performed favorably. Also, three feature groups are proved to be relative independent, and our prediction model identified the most important feature, from the collect‐and‐cart feature group. Detailed result analysis validated the effectiveness of both the extracted features and proposed machine learning methods. Copyright © 2016 John Wiley & Sons, Ltd. Yin Zhang 0006 |
Concurr. Comput. Pract. Exp. | 3 |
| 2016 | Concept over time: the combination of probabilistic topic model with wikipedia knowledge
Yin Zhang 0006, Baogang Wei, Lei Li 0005, Fei Wu 0001, Peng Zhang 0075, Yali Bian |
Expert Syst. Appl. | 2 |
| 2016 | Kernelized sparse hashing for scalable image retrieval
Yin Zhang 0006, Weiming Lu 0001, Yang Liu 0098, Fei Wu 0001 |
Neurocomputing | 1 |
| 2016 | Haptic rendering method based on generalized penetration depth computation
Yi Li 0013, Yin Zhang 0006, Xiuzi Ye, Sanyuan Zhang |
Signal Process. | 2 |
| 2015 | A question-answering system over Traditional Chinese MedicineabstractTraditional Chinese Medicine (TCM) has been around for thousands of years and it's a significant part of Chinese cultural heritage. The theoretical framework of TCM is unique and with rich of content, which contains the complex relationships between disease and medicine. Research on question-answering (QA) over TCM is significant for Chinese NLP and representative, because the resources of TCM are mostly Chinese-based. In this paper we present a QA system over TCM, which transforms user supplied questions into conjunctive query sentences (i.e. SQL) and retrieves the answer from both the built-up dataset and online encyclopedia. The contribution of this paper is threefold: Firstly, we introduce a novel approach for word segmentation over Chinese questions. We employ a TF-IDF model on the dataset to generate domain-specific dictionary with weight factor and tags, which are computed to select the best result of segmentation. Secondly, we present a novel method for constructing queries to retrieve answers. We compute the entity-attribute distance over a set of tagged words to construct incomplete ontology instances, which are used as the intermediary to generate queries. Lastly, we propose a method to integrate web data extraction with question answering, which allows us to extract answers from online encyclopedia website (i.e. Wikipedia). The results of our evaluation with 50 benchmark queries demonstrate the value of our approach. Xiangzhou Huang, Yin Zhang 0006, Baogang Wei |
BIBM | 2 |
| 2015 | Incorporating Probabilistic Knowledge into Topic Models
Yin Zhang 0006, Baogang Wei, Hongze Qian |
PAKDD (2) | 2 |
| 2015 | Topic aspect-oriented summarization via group selection
Hanyin Fang, Weiming Lu 0001, Fei Wu 0001, Yin Zhang 0006, Xindi Shang, Jian Shao 0001, Yueting Zhuang |
Neurocomputing | 4 |
| 2015 | Discovering treatment pattern in Traditional Chinese Medicine clinical cases by exploiting supervised topic model and domain knowledge
Yin Zhang 0006, Baogang Wei, Yuejiao Zhang, Xiaolin Ren, Yali Bian |
J. Biomed. Informatics | 2 |
| 2015 | Detection of engineering vehicles in high-resolution monitoring imagesabstractThis paper presents a novel formulation for detecting objects with articulated rigid bodies from high-resolution monitoring images, particularly engineering vehicles. There are many pixels in high-resolution monitoring images, and most of them represent the background. Our method first detects object patches from monitoring images using a coarse detection process. In this phase, we build a descriptor based on histograms of oriented gradient, which contain color frequency information. Then we use a linear support vector machine to rapidly detect many image patches that may contain object parts, with a low false negative rate and a high false positive rate. In the second phase, we apply a refinement classification to determine the patches that actually contain objects. In this stage, we increase the size of the image patches so that they include the complete object using models of the object parts. Then an accelerated and improved salient mask is used to improve the performance of the dense scale-invariant feature transform descriptor. The detection process returns the absolute position of positive objects in the original images. We have applied our methods to three datasets to demonstrate their effectiveness. Yin Zhang 0006, Sanyuan Zhang, Zhong-yan Liang, Xiuzi Ye |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2015 | A method for text line detection in natural images
Baogang Wei, Yonghuai Liu, Yin Zhang 0006 |
Multim. Tools Appl. | 4 |
| 2015 | The classification of multi-modal data with hidden conditional random field
Xinyang Jiang, Fei Wu 0001, Yin Zhang 0006, Siliang Tang, Weiming Lu 0001, Yueting Zhuang |
Pattern Recognit. Lett. | 3 |
| 2015 | An optimization method for penalty-based six-degrees-of-freedom haptic rendering system
Yi Li 0013, Yin Zhang 0006, Xiuzi Ye, Sanyuan Zhang |
Signal Process. Image Commun. | 2 |
| 2015 | Probabilistic Word Selection via Topic ModelingabstractWe propose selective supervised Latent Dirichlet Allocation (ssLDA) to boost the prediction performance of the widely studied supervised probabilistic topic models. We introduce a Bernoulli distribution for each word in one given document to selectthis word as a strongly or weakly discriminative one with respect to its assigned topic. The Bernoulli distribution is parameterized by the discrimination power of the word for its assigned topic. As a result, the document is represented as a “bag-of-selective-words” instead of the probabilistic “bag-of-topics” in the topic modeling domain or the flat “bag-of-words” in the traditional natural language processing domain to form a new perspective. Inheriting the general framework of supervised LDA (sLDA), ssLDA can also predict many types of response specified by a Gaussian Linear Model (GLM). Focusing on the utilization of this word selection mechanism for singe-label document classification in this paper, we conduct the variational inference for approximating the intractable posterior and derive a maximum-likelihood estimation of parameters in ssLDA. The experiments reported on textual documents show that ssLDA not only performs competitively over “state-of-the-art” classification approaches based on both the flat “bag-of-words” and probabilistic “bag-of-topics” representation in terms of classification performance, but also has the ability to discover the discrimination power of the words specified in the topics (compatible with our rational knowledge). Yueting Zhuang, Haidong Gao, Fei Wu 0001, Siliang Tang, Yin Zhang 0006, Zhongfei Zhang |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2014 | Discovering treatment pattern in traditional Chinese medicine clinical cases using topic model and domain knowledgeabstractIn Traditional Chinese Medicine (TCM), the prescription is the crystallization of clinical experience of doctors, which is the main way to cure diseases in China for thousands of years. Clinical cases, on the other hand, describe how doctors diagnose and prescribe a prescription. In this paper, we propose a framework which mines the treatment pattern in TCM clinical cases by using probabilistic topic model and TCM domain knowledge. The framework can reflect principle rules in TCM and improve function prediction of a new prescription. We evaluate our model on real world TCM clinical cases. The experiment validates the effectiveness of our method. Yin Zhang 0006, Baogang Wei, Yuejiao Zhang, Xiaolin Ren |
BIBM | 2 |
| 2014 | Cross-media hashing with kernel regressionabstractCross-media retrieval is a challenging problem in multimedia retrieval area. In the real-world, many applications involve multi-modal data, e.g., web pages containing both images and texts. How to utilize the intrinsic intra-modality and inter-modality similarity to learn the appropriate relationships of the data objects and provide efficient search across different modalities is the core of cross-media retrieval. Inspired by the fact that hashing methods well address the fast retrieval problem in the large-scale data settings, designing a cross-media hashing approach which can perform efficient retrieval over heterogenous high-dimensional feature spaces is highly desirable. In this paper, we propose a cross-media hashing approach based on kernel regression (abbreviated as KRCMH) to obtain the hash codes for the data objects across different modalities. The experiments on two real-world data sets show that KRCMH achieves superior cross-media retrieval performance comparing with the state-of-the-art methods. Zhou Yu 0001, Yin Zhang 0006, Siliang Tang, Yi Yang 0001, Qi Tian 0001, Jiebo Luo 0001 |
ICME | 2 |
| 2014 | Learning Multimodal Neural Network with Ranking ExamplesabstractTo support cross-modal information retrieval, cross-modal learning to rank approaches utilize ranking examples (e.g., an example may be a text query and its corresponding ranked images) to learn appropriate ranking (similarity) function. However, the fact that each modality is represented with intrinsically different low-level features hinders these approaches from better reducing the heterogeneity-gap between the modalities and thus giving satisfactory retrieval results. In this paper, we consider learning with neural networks, from the perspective of optimizing the listwise ranking loss of the cross-modal ranking examples. The proposed model, named Cross-Modal Ranking Neural Network (CMRNN), benefits from the advance of both neural networks on learning high-level semantics and learning to rank techniques on learning ranking function, such that the learned cross-modal ranking function is implicitly embedded in the learned high-level representation for data objects with different modalities (e.g., text and imagery) to perform cross-modal retrieval directly. We compare CMRNN to existing state-of-the-art cross-modal ranking methods on two datasets and show that it achieves a better performance. Fei Wu 0001, Xi Li 0001, Yin Zhang 0006, Weiming Lu 0001, Yueting Zhuang |
ACM Multimedia | 4 |
| 2014 | Hashing with List-Wise learning to rankabstractHashing techniques have been extensively investigated to boost similarity search for large-scale high-dimensional data. Most of the existing approaches formulate the their objective as a pair-wise similarity-preserving problem. In this paper, we consider the hashing problem from the perspective of optimizing a list-wise learning to rank problem and propose an approach called List-Wise supervised Hashing (LWH). In LWH, the hash functions are optimized by employing structural SVM in order to explicitly minimize the ranking loss of the whole list-wise permutations instead of merely the point-wise or pair-wise supervision. We evaluate the performance of LWH on two real-world data sets. Experimental results demonstrate that our method obtains a significant improvement over the state-of-the-art hashing approaches due to both structural large margin and list-wise ranking pursuing in a supervised manner. Zhou Yu 0001, Fei Wu 0001, Yin Zhang 0006, Siliang Tang, Jian Shao 0001, Yueting Zhuang |
SIGIR | 3 |
| 2014 | Structured Sparse Linear Model for Social Trust Prediction
Deng Yi, Yin Zhang 0006, Baogang Wei |
WAIM | 2 |
| 2014 | A GPU-accelerated non-negative sparse latent semantic analysis algorithm for social tagging data
Yin Zhang 0006, Deng Yi, Baogang Wei, Yueting Zhuang |
Inf. Sci. | 1 |
| 2014 | Sparse Multi-Modal HashingabstractLearning hash functions across heterogenous high-dimensional features is very desirable for many applications involving multi-modal data objects. In this paper, we propose an approach to obtain the sparse codesets for the data objects across different modalities via joint multi-modal dictionary learning, which we call sparse multi-modal hashing (abbreviated as${\rm SM}^{2}{\rm H}$). In${\rm SM}^{2}{\rm H}$, both intra-modality similarity and inter-modality similarity are first modeled by a hypergraph, then multi-modal dictionaries are jointly learned by Hypergraph Laplacian sparse coding. Based on the learned dictionaries, the sparse codeset of each data object is acquired and conducted for multi-modal approximate nearest neighbor retrieval using a sensitive Jaccard metric. The experimental results show that${\rm SM}^{2}{\rm H}$outperforms other methods in terms of mAP and Percentage on two real-world data sets. Fei Wu 0001, Zhou Yu 0001, Yi Yang 0001, Siliang Tang, Yin Zhang 0006, Yueting Zhuang |
IEEE Trans. Multim. | 5 |
| 2013 | Supervised Coupled Dictionary Learning with Group Structures for Multi-modal RetrievalabstractA better similarity mapping function across heterogeneous high-dimensional features is very desirable for many applications involving multi-modal data. In this paper, we introduce coupled dictionary learning (DL) into supervised sparse coding for multi-modal (cross-media) retrieval. We call this Supervised coupled dictionary learning with group structures for Multi-Modal retrieval (SliM2). SliM2 formulates the multi-modal mapping as a constrained dictionary learning problem. By utilizing the intrinsic power of DL to deal with the heterogeneous features, SliM2 extends unimodal DL to multi-modal DL. Moreover, the label information is employed in SliM2 to discover the shared structure inside intra-modality within the same class by a mixed norm (i.e., `l1/l2`-norm). As a result, the multimodal retrieval is conducted via a set of jointly learned mapping functions across multi-modal data. The experimental results show the effectiveness of our proposed model when applied to cross-media retrieval. Yueting Zhuang, Fei Wu 0001, Yin Zhang 0006, Weiming Lu 0001 |
AAAI | 4 |
| 2012 | Supervised cross-collection topic modelingabstractNowadays, vast amounts of multimedia data can be obtained across different collections (or domains). Therefore, it poses significant challenges for the utilization of those cross-collection data, for examples, the summarization of similarities and differences of data across different domains (e.g., CNN and NYT), as well as finding visually similar images across different visual domains (e.g., photos, paintings and hand-drawn sketches). In this paper, a supervised cross-collection Latent Dirichlet Allocation (scLDA) approach is proposed to utilize the data across different collections. As a natural extension of traditional Latent Dirichlet Allocation (LDA), scLDA not only takes the structural priors of different collections into consideration, but also exploits the category information. The strength of this work lies in integrating topic modeling, cross-domain learning and supervised learning together. We conduct scLDA for comparative text mining as well as classification of news articles and images from different collections. The results suggest that our proposed scLDA can generate meaningful collection-specific topics and achieves better retrieval accuracy than other related topic models. Haidong Gao, Siliang Tang, Yin Zhang 0006, Dapeng Jiang, Fei Wu 0001, Yueting Zhuang |
ACM Multimedia | 3 |
| 2011 | A Method for Finding Groups of Related Herbs in Traditional Chinese Medicine
Yin Zhang 0006, Baogang Wei |
ADMA (1) | 2 |
| 2011 | Tag Clustering and Refinement on Semantic Unity GraphabstractRecently, there has been extensive research towards the user-provided tags on photo sharing websites which can greatly facilitate image retrieval and management. However, due to the arbitrariness of the tagging activities, these tags are often imprecise and incomplete. As a result, quite a few technologies has been proposed to improve the user experience on these photo sharing systems, including tag clustering and refinement, etc. In this work, we propose a novel framework to model the relationships among tags and images which can be applied to many tag based applications. Different from previous approaches which model images and tags as heterogeneous objects, images and their tags are uniformly viewed as compositions of Semantic Unities in our framework. Then Semantic Unity Graph (SUG) is introduced to represent the complex and high-order relationships among these Semantic Unities. Based on the representation of Semantic Unity Graph, the relevance of images and tags can be naturally measured in terms of the similarity of their Semantic Unities. Then Tag clustering and refinement can then be performed on SUG and the polysemy of images and tags is explicitly considered in this framework. The experiment results conducted on NUS-WIDE and MIR-Flickr datasets demonstrate the effectiveness and efficiency of the proposed approach. Yang Liu 0098, Fei Wu 0001, Yin Zhang 0006, Jian Shao 0001, Yueting Zhuang |
ICDM | 3 |
| 2011 | Hypergraph spectral hashing for similarity search of social imageabstractThe development of social media brings great challenges to image retrieval on both efficiency and accuracy. In addition to achieving fast similarity search over large scale data, it is very crucial to represent the complex and high-order relationships among the social contents to improve the semantic understanding of social images.In this paper, unified hypergraph is implemented to model the various relationships among images and other contexts in social media. Moreover, we extend traditional spectral hashing to hypergraph to accelerate similarity search of social images by mapping semantically related vertices into similar binary codes within a short Hamming distance. Furthermore, the proposed HSH approach is extended to out-of-sample data in a supervised manner. We evaluated our approach on the dataset crawled from Flickr and the experiment results indicate that our proposed HSH approach is both efficient and effective. Yueting Zhuang, Yang Liu 0098, Fei Wu 0001, Yin Zhang 0006, Jian Shao 0001 |
ACM Multimedia | 4 |
| 2010 | Javelin: an access and manipulation interface for large displaysabstractWe describe a user interface and interaction technique, named ‘Javelin’, designed for large display environments. It provides quick access to random screen regions and manipulation methods for screen widgets which are difficult or impossible to reach. It consists of a dynamic global thumbnail, a touchpad widget that drives the screen cursor, and a teleport widget in which interactions are transferred to its target screen region. Javelin can be easily integrated into many programs to optimize their interaction performance in large screens. The experiment and user study show that Javelin can extend user access field and enhance widget manipulation in large displays. Zhenkun Zhou, Jiangqin Wu, Yin Zhang 0006, Da-wei Xie, Yueting Zhuang |
J. Zhejiang Univ. Sci. C | 3 |
| 2008 | RNACompress: Grammar-based compression and informational complexity measurement of RNA secondary structureabstractBACKGROUND: With the rapid emergence of RNA databases and newly identified non-coding RNAs, an efficient compression algorithm for RNA sequence and structural information is needed for the storage and analysis of such data. Although several algorithms for compressing DNA sequences have been proposed, none of them are suitable for the compression of RNA sequences with their secondary structures simultaneously. This kind of compression not only facilitates the maintenance of RNA data, but also supplies a novel way to measure the informational complexity of RNA structural data, raising the possibility of studying the relationship between the functional activities of RNA structures and their complexities, as well as various structural properties of RNA based on compression. RESULTS: RNACompress employs an efficient grammar-based model to compress RNA sequences and their secondary structures. The main goals of this algorithm are two fold: (1) present a robust and effective way for RNA structural data compression; (2) design a suitable model to represent RNA secondary structure as well as derive the informational complexity of the structural data based on compression. Our extensive tests have shown that RNACompress achieves a universally better compression ratio compared with other sequence-specific or common text-specific compression algorithms, such as Gencompress, winrar and gzip. Moreover, a test of the activities of distinct GTP-binding RNAs (aptamers) compared with their structural complexity shows that our defined informational complexity can be used to describe how complexity varies with activity. These results lead to an objective means of comparing the functional properties of heteropolymers from the information perspective. CONCLUSION: A universal algorithm for the compression of RNA secondary structure as well as the evaluation of its informational complexity is discussed in this paper. We have developed RNACompress, as a useful tool for academic users. Extensive tests have shown that RNACompress is a universally efficient algorithm for the compression of RNA sequences with their secondary structures. RNACompress also serves as a good measurement of the informational complexity of RNA secondary structure, which can be used to study the functional activities of RNA molecules. Chun Chen 0001, Jiajun Bu, Yin Zhang 0006, Xiuzi Ye |
BMC Bioinform. | 5 |
| 2005 | Uniform color transferabstractWe present in this paper a general algorithm called uniform color transfer for transferring dominant colors in one image to another. Our algorithm frees users from interactions required for colorizing grayscale image or transferring color between images. Different from existing algorithms, we cluster the source color image into regions using the GMM-EM method, and the target image using the K-means algorithm. We then impose chromatic mean values on corresponding source image regions in the target image. Experiments showed that our algorithm can give better results. Shuchang Xu, Yin Zhang 0006, Sanyuan Zhang, Xiuzi Ye |
ICIP (3) | 2 |
| 2005 | Radius-Normal Histogram and Hybrid Strategy for 3D Shape RetrievalabstractRecent development in computer hardware and modeling technologies has led to a fast increasing of 3D models. To help user find expected models accurately and efficiently, technologies on content-based 3D model retrievals have become one of the most challenging research topic recently. In this paper, radius-normal histogram(RNH) is proposed to describe shape contents and used for shape retrievals. The RNH shape descriptor first uses a series of concentric spheres to capture the point distribution information of the given model. Then for points in each concentric sphere, a radian normal angle is computed to extract the local geometry features. Finally, the radius-normal histogram is constructed by using the extracted shape signatures. The proposed shape representation remains invariant under rotations. It can be generated from the given 3D model efficiently and easily as well. Performance comparisons for the shape benchmark database have proven that the proposed algorithm can achieve better retrieving performance than other similar histogram-based shape representations. To further improve the retrieving performance, different shape features are combined based on dynamic weight selections and a hierarchical architecture is used to speed up the matching process. Yin Zhang 0006, Sanyuan Zhang, Xiuzi Ye |
SMI | 2 |