EDBT 2026 Demo / reviewers in the wild / expert
Hao Zhang 0048
dblp:55/2270-48
· DBLP profile ↗
33ranked-venue papers
4as first author
27since 2021 · last 2026
0000-0002-2725-6458ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 3 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hesitation and Tolerance in Recommender SystemsabstractUsers’ interactions with recommender systems often involve more than simple acceptance or rejection. We highlight two overlooked states: hesitation, when people deliberate without certainty, and tolerance, when this hesitation escalates into unwanted engagement before ending in disinterest. Across two large-scale surveys (N = 6, 644 and N = 3, 864), hesitation was nearly universal, and tolerance emerged as a recurring source of wasted time, frustration, and diminished trust. Analyses of e-commerce and short-video platforms confirm that tolerance behaviors, such as clicking without purchase or shallow viewing, correlate with decreased activity. Finally, an online field study at scale shows that even lightweight strategies treating tolerance as distinct from interest can improve retention while reducing wasted effort. By surfacing hesitation and tolerance as consequential states, this work reframes how recommender systems should interpret feedback, moving beyond clicks and dwell time toward designs that respect user value, reduce hidden costs, and sustain engagement. Kuan Zou, Aixin Sun, Yitong Ji, Hao Zhang 0048, Jing Wang 0060, Zhuohao (Jerry) Zhang, Xuemeng Jiang |
CHI | 4 |
| 2025 | FineReason: Evaluating and Improving LLMs' Deliberate Reasoning through Reflective Puzzle SolvingabstractGuizhen Chen, Weiwen Xu, Hao Zhang, Hou Pong Chan, Chaoqun Liu, Lidong Bing, Deli Zhao, Anh Tuan Luu, Yu Rong. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Guizhen Chen, Weiwen Xu, Hao Zhang 0048, Hou Pong Chan, Chaoqun Liu, Lidong Bing, Deli Zhao, Anh Tuan Luu, Yu Rong 0001 |
ACL (1) | 3 |
| 2025 | CoIR: A Comprehensive Benchmark for Code Information Retrieval ModelsabstractXiangyang Li, Kuicai Dong, Yi Quan Lee, Wei Xia, Hao Zhang, Xinyi Dai, Yasheng Wang, Ruiming Tang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xiangyang Li 0004, Kuicai Dong, Yi Quan Lee, Wei Xia 0001, Hao Zhang 0048, Xinyi Dai, Yasheng Wang, Ruiming Tang |
ACL (1) | 5 |
| 2025 | Adaptive Tool Use in Large Language Models with Meta-Cognition TriggerabstractWenjun Li, Dexun Li, Kuicai Dong, Cong Zhang, Hao Zhang, Weiwen Liu, Yasheng Wang, Ruiming Tang, Yong Liu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Dexun Li, Kuicai Dong, Hao Zhang 0048, Weiwen Liu, Yasheng Wang, Ruiming Tang, Yong Liu 0020 |
ACL (1) | 5 |
| 2025 | Analyzing LLMs' Knowledge Boundary Cognition Across Languages Through the Lens of Internal RepresentationsabstractWhile understanding the knowledge boundaries of LLMs is crucial to prevent hallucination, research on the knowledge boundaries of LLMs has predominantly focused on English. In this work, we present the first study to analyze how LLMs recognize knowledge boundaries across different languages by probing their internal representations when processing known and unknown questions in multiple languages. Our empirical studies reveal three key findings: 1) LLMs' perceptions of knowledge boundaries are encoded in the middle to middle-upper layers across different languages. 2) Language differences in knowledge boundary perception follow a linear structure, which motivates our proposal of a training-free alignment method that effectively transfers knowledge boundary perception ability across languages, thereby helping reduce hallucination risk in low-resource languages; 3) Fine-tuning on bilingual question pair translation further enhances LLMs' recognition of knowledge boundaries across languages. Given the absence of standard testbeds for cross-lingual knowledge boundary analysis, we construct a multilingual evaluation suite comprising three representative types of knowledge boundary data. Our code and datasets are publicly available at https://github.com/DAMO-NLP-SG/ LLM-Multilingual-Knowledge-Boundaries. Chenghao Xiao, Hou Pong Chan, Hao Zhang 0048, Mahani Aljunied, Lidong Bing, Noura Al Moubayed, Yu Rong 0001 |
ACL (1) | 3 |
| 2025 | ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical ReasoningabstractYu Sun, Xingyu Qian, Weiwen Xu, Hao Zhang, Chenghao Xiao, Long Li, Deli Zhao, Wenbing Huang, Tingyang Xu, Qifeng Bai, Yu Rong. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Xingyu Qian, Weiwen Xu, Hao Zhang 0048, Chenghao Xiao, Deli Zhao, Wenbing Huang 0001, Tingyang Xu, Qifeng Bai, Yu Rong 0001 |
EMNLP | 4 |
| 2025 | SafetyQuizzer: Timely and Dynamic Evaluation on the Safety of LLMsabstractZhichao Shi, Shaoling Jing, Yi Cheng, Hao Zhang, Yuanzhuo Wang, Jie Zhang, Huawei Shen, Xueqi Cheng. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Zhichao Shi 0001, Shaoling Jing, Hao Zhang 0048, Yuanzhuo Wang, Jie Zhang 0071, Huawei Shen, Xueqi Cheng 0001 |
NAACL (Long Papers) | 4 |
| 2025 | Scaling Language-centric Omnimodal Representation LearningabstractRecent multimodal embedding approaches leveraging multimodal large language models (MLLMs) fine-tuned with contrastive learning (CL) have shown promising results, yet the underlying reasons behind their superiority remain underexplored. This work argues that a crucial advantage of MLLM-based approaches stems from implicit cross-modal alignment achieved during generative pretraining, where the language decoder learns to exploit multimodal signals within a shared representation space for generating unimodal outputs. Through analysis of anisotropy and kernel similarity structure, we empirically confirm that latent alignment emerges within MLLM representations, allowing CL to serve as a lightweight refinement stage. Leveraging this insight, we propose a Language-Centric Omnimodal Embedding framework, termed LCO-Embed. Extensive experiments across diverse backbones and benchmarks demonstrate its effectiveness, achieving state-of-the-art performance across modalities. Furthermore, we identify a Generation-Representation Scaling Law (GRSL), showing that the representational capabilities gained through contrastive refinement scale positively with the MLLM's generative capabilities. This suggests that improving generative abilities evolves as an effective paradigm for enhancing representation quality. We provide a theoretical explanation of GRSL, which formally links the MLLM's generative quality to the upper bound on its representation performance, and validate it on a challenging, low-resource visual-document retrieval task, showing that continual generative pretraining before CL can further enhance the potential of a model's embedding capabilities. Codes, models, and resources are available at https://github.com/LCO-Embedding/LCO-Embedding. Chenghao Xiao, Hou Pong Chan, Hao Zhang 0048, Weiwen Xu, Mahani Aljunied, Yu Rong 0001 |
NeurIPS | 3 |
| 2025 | Revisiting the Design of In-Memory Dynamic Graph StorageabstractThe effectiveness of in-memory dynamic graph storage (DGS) for supporting concurrent graph read and write queries is crucial for real-time graph analytics and updates. Various methods have been proposed, for example, LLAMA, Aspen, LiveGraph, Teseo, and Sortledton. These approaches differ significantly in their support for read and write operations, space overhead, and concurrency control. However, there has been no systematic study to explore the trade-offs among these dimensions. In this paper, we evaluate the effectiveness of individual techniques and identify the performance factors affecting these storage methods by proposing a common abstraction for DGS design and implementing a generic test framework based on this abstraction. Our findings highlight several key insights: 1) Existing DGS methods exhibit substantial space overhead. For example, Aspen consumes 3.3-10.8x more memory than CSR, while the optimal fine-grained methods consume 4.1-8.9x more memory than CSR, indicating a significant memory overhead. 2) Existing methods often overlook memory access impact of modern architectures, leading to performance degradation compared to continuous storage methods. 3) Fine-grained concurrency control methods, in particular, suffer from severe efficiency and space issues due to maintaining versions and performing checks for each neighbor. These methods also experience significant contention on high-degree vertices. Our systematic study reveals these performance bottlenecks and outlines future directions to improve DGS for real-time graph analytics. Jixian Su, Chiyu Hao, Shixuan Sun, Hao Zhang 0048, Yao Chen 0008, Chenyi Zhang 0002, Bingsheng He, Minyi Guo |
Proc. ACM Manag. Data | 4 |
| 2025 | How Can Recommender Systems Benefit from Large Language Models: A SurveyabstractWith the rapid development of online services and web applications, recommender systems (RS) have become increasingly indispensable for mitigating information overload and matching users’ information needs by providing personalized suggestions over items. Although the RS research community has made remarkable progress over the past decades, conventional recommendation models (CRM) still have some limitations, e.g., lacking open-domain world knowledge, and difficulties in comprehending users’ underlying preferences and motivations. Meanwhile, large language models (LLM) have shown impressive general intelligence and human-like capabilities for various natural language processing (NLP) tasks, which mainly stem from their extensive open-world knowledge, logical and commonsense reasoning abilities, as well as their comprehension of human culture and society. Consequently, the emergence of LLM is inspiring the design of RS and pointing out a promising research direction, i.e., whether we can incorporate LLM and benefit from their common knowledge and capabilities to compensate for the limitations of CRM. In this article, we conduct a comprehensive survey on this research direction, and draw a bird’s-eye view from the perspective of the whole pipeline in real-world RS. Specifically, we summarize existing research works from two orthogonal aspects: where and how to adapt LLM to RS. For the “ WHERE ” question, we discuss the roles that LLM could play in different stages of the recommendation pipeline, i.e., feature engineering, feature encoder, scoring/ranking function, user interaction, and pipeline controller. For the “ HOW ” question, we investigate the training and inference strategies, resulting in two fine-grained taxonomy criteria, i.e., whether to tune LLM or not during training, and whether to involve CRM for inference. Detailed analysis and general development paths are provided for both “WHERE” and “HOW” questions, respectively. Then, we highlight the key challenges in adapting LLM to RS from three aspects, i.e., efficiency, effectiveness, and ethics. Finally, we summarize the survey and discuss the future prospects. Jianghao Lin, Xinyi Dai, Yunjia Xi, Weiwen Liu, Bo Chen 0023, Hao Zhang 0048, Yong Liu 0020, Chuhan Wu, Xiangyang Li 0004, Chenxu Zhu, Huifeng Guo, Yong Yu 0001, Ruiming Tang, Weinan Zhang 0001 |
ACM Trans. Inf. Syst. | 6 |
| 2024 | Collaborative Cross-modal Fusion with Large Language Model for RecommendationabstractDespite the success of conventional collaborative filtering (CF) approaches for recommendation systems, they exhibit limitations in leveraging semantic knowledge within the textual attributes of users and items. Recent focus on the application of large language models for recommendation (LLM4Rec) has highlighted their capability for effective semantic knowledge capture. However, these methods often overlook the collaborative signals in user behaviors. Some simply instruct-tune a language model, while others directly inject the embeddings of a CF-based model, lacking a synergistic fusion of different modalities. To address these issues, we propose a framework of Collaborative Cross-modal Fusion with Large Language Models, termed CCF-LLM, for recommendation. In this framework, we translate the user-item interactions into a hybrid prompt to encode both semantic knowledge and collaborative signals, and then employ an attentive cross-modal fusion strategy to effectively fuse latent embeddings of both modalities. Extensive experiments demonstrate that CCF-LLM outperforms existing methods by effectively utilizing semantic and collaborative signals in the LLM4Rec context. Zhongzhou Liu, Hao Zhang 0048, Kuicai Dong, Yuan Fang 0001 |
CIKM | 2 |
| 2024 | Retrieval-Oriented Knowledge for Click-Through Rate PredictionabstractClick-through rate (CTR) prediction is crucial for personalized online services. Sample-level retrieval-based models, such as RIM, have demonstrated remarkable performance. However, they face challenges including inference inefficiency and high resource consumption due to the retrieval process, which hinder their practical application in industrial settings. To address this, we propose a universal plug-and-play retrieval-oriented knowledge (ROK) framework that bypasses the real retrieval process. The framework features a knowledge base that preserves and imitates the retrieved & aggregated representations using a decomposition-reconstruction paradigm. Knowledge distillation and contrastive learning optimize the knowledge base, enabling the integration of retrieval-enhanced representations with various CTR models. Experiments on three large-scale datasets demonstrate ROK's exceptional compatibility and performance, with the neural knowledge base serving as an effective surrogate for the retrieval pool. ROK surpasses the teacher model while maintaining superior inference efficiency and demonstrates the feasibility of distilling knowledge from non-parametric methods using a parametric approach. These results highlight ROK's strong potential for real-world applications and its ability to transform retrieval-based methods into practical solutions. Our implementation code is available to support reproducibility1. Huanshuo Liu, Bo Chen 0023, Menghui Zhu, Jianghao Lin, Jiarui Qin, Hao Zhang 0048, Yang Yang 0001, Ruiming Tang |
CIKM | 6 |
| 2024 | Parameter-Efficient Conversational Recommender System as a Language Processing TaskabstractConversational recommender systems (CRS) aim to recommend relevant items to users by eliciting user preference through natural language conversation.Prior work often utilizes external knowledge graphs for items' semantic information, a language model for dialogue generation, and a recommendation module for ranking relevant items.This combination of multiple components suffers from a cumbersome training process, and leads to semantic misalignment issues between dialogue generation and item recommendation.In this paper, we represent items in natural language and formulate CRS as a natural language processing task.Accordingly, we leverage the power of pre-trained language models to encode items, understand user intent via conversation, perform item recommendation through semantic matching, and generate dialogues.As a unified model, our PECRS (Parameter-Efficient CRS), can be optimized in a single stage, without relying on non-textual metadata such as a knowledge graph.Experiments on two benchmark CRS datasets, ReDial and INSPIRED, demonstrate the effectiveness of PECRS on recommendation and conversation.Our Mathieu Ravaut, Hao Zhang 0048, Aixin Sun, Yong Liu 0020 |
EACL (1) | 2 |
| 2024 | Learning Feature Semantic Matching for Spatio-Temporal Video GroundingabstractSpatio-temporal video grounding (STVG) aims to localize a spatio-temporal tube, including temporal boundaries and object bounding boxes, that semantically corresponds to a given language description in an untrimmed video. The existing onestage solutions in this task face two significant challenges, namely, vision-text semantic misalignment and spatial mislocalization, which limit their performance in grounding. These two limitations are mainly caused by neglect of fine-grained alignment in crossmodality fusion and the reliance on a text-agnostic query in sequentially spatial localization. To address these issues, we propose an effective model with a newly designed Feature Semantic Matching (FSM) module based on a Transformer architecture to address the above issues. Our method introduces a crossmodal feature matching module to achieve multi-granularity alignment between video and text while preventing the weakening of important features during the feature fusion stage. Additionally, we design a query-modulated matching module to facilitate text-relevant tube construction by multiple query generation and tubulet sequence matching. To ensure the quality of tube construction, we employ a novel mismatching rectify contrastive loss to rectify the mismatching between the learnable query and the objects corresponding to the text descriptions by restricting the generated spatial query. Extensive experiments demonstrate that our method outperforms the state-of-the-art methods on two challenging STVG benchmarks. Hao Fang 0010, Hao Zhang 0048, Jialin Gao, Xiankai Lu, Xiushan Nie, Yilong Yin |
IEEE Trans. Multim. | 3 |
| 2024 | Relational Network via Cascade CRF for Video Language GroundingabstractVideo Language Grounding is one of the most challenging cross-modal video understanding tasks. This task aims to localize a target moment semantically corresponding to a given language query in an untrimmed video. Many existing VLG methods rely on the proposal-based framework, despite the dominant performance achieved, they usually focus on interacting a few internal frames with the query to score segment proposals, trapping in the long-range dependencies when the proposal feature is limited. Meanwhile, adjacent proposals share similar visual semantics, making VLG models hard to align the accurate semantics of video-query contents and degenerating the ranking performance. To remedy the above limitations, we propose VLG-CRF by introducing the conditional random fields (CRFs) to handle the discrete yet indistinguishable proposals. Specifically, VLG-CRF consists of two cascade CRF-based modules. The AttentiveCRFs is developed for multi-modal feature fusion to better integrate temporal and semantic relation between modalities. We also devise a new variant of ConvCRFs to capture the relation of discrete segments and rectify the predicting scores to make relatively high prediction scores clustered in a range. Experiments on three benchmark datasets,i.e., Charades-STA, ActivityNet-Caption, and TACoS, show the superiority of our method and the state-of-the-art performance is achieved. Xiankai Lu, Hao Zhang 0048, Xiushan Nie, Yilong Yin, Jianbing Shen |
IEEE Trans. Multim. | 3 |
| 2023 | MS-DETR: Natural Language Video Localization with Sampling Moment-Moment InteractionabstractGiven a query, the task of Natural Language Video Localization (NLVL) is to localize a temporal moment in an untrimmed video that semantically matches the query.In this paper, we adopt a proposal-based solution that generates proposals (i.e., candidate moments) and then select the best matching proposal.On top of modeling the cross-modal interaction between candidate moments and the query, our proposed Moment Sampling DETR (MS-DETR) enables efficient moment-moment relation modeling.The core idea is to sample a subset of moments guided by the learnable templates with an adopted DETR (DEtection TRansformer) framework.To achieve this, we design a multiscale visual-linguistic encoder, and an anchorguided moment decoder paired with a set of learnable templates.Experimental results on three public datasets demonstrate the superior performance of MS-DETR. 1 Jing Wang 0060, Aixin Sun, Hao Zhang 0048, Xiaoli Li 0001 |
ACL (1) | 3 |
| 2023 | WSDM 2023 Workshop on Interactive Recommender SystemsabstractInteractive recommender systems have attracted increasingly research attentions from both academia and industry. This workshop is a half-day event, which provides a forum for researchers and practitioners to discuss recent research progress and novel research directions about interactive recommender systems. The program will include two keynotes and 6 to 8 research paper presentations. The objective of this workshop is to consolidate the recent technical progresses about interactive recommendation, which will be a promising research and development direction for future recommendation technologies. This workshop will attract the attention of researchers from both academia and industry. It aligns with WSDM's spirit of promoting the collaborations between academia and industry. Yong Liu 0020, Hao Zhang 0048, Zhu Sun 0001, Shoujin Wang, Jie Zhang 0002 |
WSDM | 2 |
| 2023 | Temporal Sentence Grounding in Videos: A Survey and Future DirectionsabstractTemporal sentence grounding in videos (TSGV), a.k.a., natural language video localization (NLVL) or video moment retrieval (VMR), aims to retrieve a temporal moment that semantically corresponds to a language query from an untrimmed video. Connecting computer vision and natural language, TSGV has drawn significant attention from researchers in both communities. This survey attempts to provide a summary of fundamental concepts in TSGV and current research status, as well as future research directions. As the background, we present a common structure of functional components in TSGV, in a tutorial style: from feature extraction from raw video and language query, to answer prediction of the target moment. Then we review the techniques for multimodal understanding and interaction, which is the key focus of TSGV for effective alignment between the two modalities. We construct a taxonomy of TSGV techniques and elaborate the methods in different categories with their strengths and weaknesses. Lastly, we discuss issues with the current TSGV research and share our insights about promising research directions. Hao Zhang 0048, Aixin Sun, Joey Tianyi Zhou |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Interventional Training for Out-Of-Distribution Natural Language UnderstandingabstractOut-of-distribution (OOD) settings are used to measure a model's performance when the distribution of the test data is different from that of the training data.NLU models are known to suffer in OOD settings (Utama et al., 2020b).We study this issue from the perspective of causality, which sees confounding bias as the reason for models to learn spurious correlations.While a common solution is to perform intervention, existing methods handle only known and single confounder (Pearl and Mackenzie, 2018), but in many NLU tasks the confounders can be both unknown and multifactorial.In this paper, we propose a novel interventional training method called Bottom-up Automatic Intervention (BAI) that performs multi-granular intervention with identified multifactorial confounders.Our experiments on three NLU tasks, namely, natural language inference, fact verification and paraphrase identification, show the effectiveness of BAI for tackling different OOD settings.1 Sicheng Yu, Jing Jiang 0001, Hao Zhang 0048, Yulei Niu, Qianru Sun, Lidong Bing |
EMNLP | 3 |
| 2022 | An Embarrassingly Simple Model for Dialogue Relation ExtractionabstractDialogue relation extraction (RE) is to predict the relation type of two entities mentioned in a dialogue. In this paper, we propose a simple yet effective model named SimpleRE for the RE task. SimpleRE captures the interrelations among multiple relations in a dialogue through a novel input format named BERT Relation Token Sequence (BRS). In BRS, multiple [CLS] tokens are used to capture possible relations between different pairs of entities mentioned in the dialogue. A Relation Refinement Gate (RRG) is then designed to extract relation-specific semantic representation in an adaptive manner. Experiments on the DialogRE dataset show that SimpleRE achieves the best performance, with much shorter training time. Further, SimpleRE outperforms all direct baselines on sentence-level RE without using external resources. Fuzhao Xue, Aixin Sun, Hao Zhang 0048, Jinjie Ni, Chng Eng Siong |
ICASSP | 3 |
| 2022 | Context Modeling with Evidence Filter for Multiple Choice Question AnsweringabstractMultiple-Choice Question Answering (MCQA) is one of the challenging tasks in machine reading comprehension. The main challenge in MCQA is to extract "evidence" from the given context that supports the correct answer. In OpenbookQA dataset [1], the requirement of extracting "evidence" is particularly important due to the mutual independence of sentences in the context. Existing work tackles this problem by annotated evidence or distant supervision with rules which overly rely on human efforts. To address the challenge, we propose a simple yet effective approach termed evidence filtering to model the relationships between the encoded contexts with respect to different options collectively, and to potentially highlight the evidence sentences and filter out unrelated sentences. In addition to the effective reduction of human efforts of our approach compared, through extensive experiments on OpenbookQA, we show that the proposed approach outperforms the models that use the same backbone and more training data; and our parameter analysis also demonstrates the interpretability of our approach. Sicheng Yu, Hao Zhang 0048, Jing Jiang 0001 |
ICASSP | 2 |
| 2022 | Natural Language Video Localization: A Revisit in Span-Based Question Answering FrameworkabstractNatural Language Video Localization (NLVL) aims to locate a target moment from an untrimmed video that semantically corresponds to a text query. Existing approaches mainly solve the NLVL problem from the perspective of computer vision by formulating it as ranking, anchor, or regression tasks. These methods suffer from large performance degradation when localizing on long videos. In this work, we address the NLVL from a new perspective, i.e., span-based question answering (QA), by treating the input video as a text passage. We propose a video span localizing network (VSLNet), on top of the standard span-based QA framework (named VSLBase), to address NLVL. VSLNet tackles the differences between NLVL and span-based QA through a simple yet effective query-guided highlighting (QGH) strategy. QGH guides VSLNet to search for the matching video span within a highlighted region. To address the performance degradation on long videos, we further extend VSLNet to VSLNet-L by applying a multi-scale split-and-concatenation strategy. VSLNet-L first splits the untrimmed video into short clip segments; then, it predicts which clip segment contains the target moment and suppresses the importance of other segments. Finally, the clip segments are concatenated, with different confidences, to locate the target moment accurately. Extensive experiments on three benchmark datasets show that the proposed VSLNet and VSLNet-L outperform the state-of-the-art methods; VSLNet-L addresses the issue of performance degradation on long videos. Our study suggests that the span-based QA framework is an effective strategy to solve the NLVL problem. Hao Zhang 0048, Aixin Sun, Liangli Zhen, Joey Tianyi Zhou, Rick Siow Mong Goh |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | GDPNet: Refining Latent Multi-View Graph for Relation ExtractionabstractRelation Extraction (RE) is to predict the relation type of two entities that are mentioned in a piece of text, e.g., a sentence or a dialogue. When the given text is long, it is challenging to identify indicative words for the relation prediction. Recent advances on RE task are from BERT-based sequence modeling and graph-based modeling of relationships among the tokens in the sequence. In this paper, we propose to construct a latent multi-view graph to capture various possible relationships among tokens. We then refine this graph to select important words for relation prediction. Finally, the representation of the refined graph and the BERT-based sequence representation are concatenated for relation extraction. Specifically, in our proposed GDPNet (Gaussian Dynamic Time Warping Pooling Net), we utilize Gaussian Graph Generator (GGG) to generate edges of the multi-view graph. The graph is then refined by Dynamic Time Warping Pooling (DTWPool). On DialogRE and TACRED, we show that GDPNet achieves the best performance on dialogue-level RE, and comparable performance with the state-of-the-arts on sentence-level RE. Our code is available at https://github.com/XueFuzhao/GDPNet. Fuzhao Xue, Aixin Sun, Hao Zhang 0048, Chng Eng Siong |
AAAI | 3 |
| 2021 | COSY: COunterfactual SYntax for Cross-Lingual UnderstandingabstractSicheng Yu, Hao Zhang, Yulei Niu, Qianru Sun, Jing Jiang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Sicheng Yu, Hao Zhang 0048, Yulei Niu, Qianru Sun, Jing Jiang 0001 |
ACL/IJCNLP (1) | 2 |
| 2021 | Interventional Video Grounding With Dual Contrastive LearningabstractVideo grounding aims to localize a moment from an untrimmed video for a given textual query. Existing approaches focus more on the alignment of visual and language stimuli with various likelihood-based matching or regression strategies, i.e., P(Y |X). Consequently, these models may suffer from spurious correlations between the language and video features due to the selection bias of the dataset. 1) To uncover the causality behind the model and data, we first propose a novel paradigm from the perspective of the causal inference, i.e., interventional video grounding (IVG) that leverages backdoor adjustment to deconfound the selection bias based on structured causal model (SCM) and do-calculus P(Y |do(X)). Then, we present a simple yet effective method to approximate the unobserved confounder as it cannot be directly sampled from the dataset. 2) Meanwhile, we introduce a dual contrastive learning approach (DCL) to better align the text and video by maximizing the mutual information (MI) between query and video clips, and the MI between start/end frames of a target moment and the others within a video to learn more informative visual representations. Experiments on three standard benchmarks show the effectiveness of our approaches. Guoshun Nan, Rui Qiao 0006, Jun Liu 0036, Sicong Leng, Hao Zhang 0048, Wei Lu 0011 |
CVPR | 6 |
| 2021 | Video Corpus Moment Retrieval with Contrastive LearningabstractGiven a collection of untrimmed and unsegmented videos, video corpus moment retrieval (VCMR) is to retrieve a temporal moment (i.e., a fraction of a video) that semantically corresponds to a given text query. As video and text are from two distinct feature spaces, there are two general approaches to address VCMR: (i) to separately encode each modality representations, then align the two modality representations for query processing, and (ii) to adopt fine-grained cross-modal interaction to learn multi-modal representations for query processing. While the second approach often leads to better retrieval accuracy, the first approach is far more efficient. In this paper, we propose a Retrieval and Localization Network with Contrastive Learning (ReLoCLNet) for VCMR. We adopt the first approach and introduce two contrastive learning objectives to refine video encoder and text encoder to learn video and text representations separately but with better alignment for VCMR. The video contrastive learning (VideoCL) is to maximize mutual information between query and candidate video at video-level. The frame contrastive learning (FrameCL) aims to highlight the moment region corresponds to the query at frame-level, within a video. Experimental results show that, although ReLoCLNet encodes text and video separately for efficiency, its retrieval accuracy is comparable with baselines adopting cross-modal interaction learning. Hao Zhang 0048, Aixin Sun, Guoshun Nan, Liangli Zhen, Joey Tianyi Zhou, Rick Siow Mong Goh |
SIGIR | 1 |
| 2021 | Dual Adversarial Transfer for Sequence LabelingabstractWe propose a new architecture for addressing sequence labeling, termed Dual Adversarial Transfer Network (DATNet). Specifically, the proposed DATNet includes two variants, i.e., DATNet-F and DATNet-P, which are proposed to explore effective feature fusion between high and low resource. To address the noisy and imbalanced training data, we propose a novel Generalized Resource-Adversarial Discriminator (GRAD) and adopt adversarial training to boost model generalization. We investigate the effects of different components of DATNet across different domains and languages, and show that significant improvement can be obtained especially for low-resource data. Without augmenting any additional hand-crafted features, we achieve state-of-the-art performances on CoNLL, Twitter, PTB-WSJ, OntoNotes and Universal Dependencies with three popular sequence labeling tasks, i.e., Named entity recognition (NER), Part-of-Speech (POS) Tagging and Chunking. Joey Tianyi Zhou, Hao Zhang 0048, Di Jin 0005, Xi Peng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | RoboCoDraw: Robotic Avatar Drawing with GAN-Based Style Transfer and Time-Efficient Path OptimizationabstractRobotic drawing has become increasingly popular as an entertainment and interactive tool. In this paper we present RoboCoDraw, a real-time collaborative robot-based drawing system that draws stylized human face sketches interactively in front of human users, by using the Generative Adversarial Network (GAN)-based style transfer and a Random-Key Genetic Algorithm (RKGA)-based path optimization. The proposed RoboCoDraw system takes a real human face image as input, converts it to a stylized avatar, then draws it with a robotic arm. A core component in this system is the AvatarGAN proposed by us, which generates a cartoon avatar face image from a real human face. AvatarGAN is trained with unpaired face and avatar images only and can generate avatar images of much better likeness with human face images in comparison with the vanilla CycleGAN. After the avatar image is generated, it is fed to a line extraction algorithm and converted to sketches. An RKGA-based path optimization algorithm is applied to find a time-efficient robotic drawing path to be executed by the robotic arm. We demonstrate the capability of RoboCoDraw on various face images using a lightweight, safe collaborative robot UR5. Tianying Wang, Wei Qi Toh, Hao Zhang 0048, Xiuchao Sui, Shaohua Li 0003, Yong Liu 0026 |
AAAI | 3 |
| 2020 | Multi-source Meta Transfer for Low Resource Multiple-Choice Question AnsweringabstractMultiple-choice question answering (MCQA) is one of the most challenging tasks in machine reading comprehension since it requires more advanced reading comprehension skills such as logical reasoning, summarization, and arithmetic operations.Unfortunately, most existing MCQA datasets are small in size, which increases the difficulty of model learning and generalization.To address this challenge, we propose a multi-source meta transfer (MMT) for low-resource MCQA.In this framework, we first extend meta learning by incorporating multiple training sources to learn a generalized feature representation across domains.To bridge the distribution gap between training sources and the target, we further introduce the meta transfer that can be integrated into the multi-source meta training.More importantly, the proposed MMT is independent of backbone language models.Extensive experiments demonstrate the superiority of MMT over state-of-the-arts, and continuous improvements can be achieved on different backbone networks on both supervised and unsupervised domain adaptation settings. Ming Yan 0007, Hao Zhang 0048, Di Jin 0005, Joey Tianyi Zhou |
ACL | 2 |
| 2020 | Span-based Localizing Network for Natural Language Video LocalizationabstractGiven an untrimmed video and a text query, natural language video localization (NLVL) is to locate a matching span from the video that semantically corresponds to the query.Existing solutions formulate NLVL either as a ranking task and apply multimodal matching architecture, or as a regression task to directly regress the target video span.In this work, we address NLVL task with a span-based QA approach by treating the input video as text passage.We propose a video span localizing network (VSLNet), on top of the standard span-based QA framework, to address NLVL.The proposed VSLNet tackles the differences between NLVL and span-based QA through a simple and yet effective query-guided highlighting (QGH) strategy.The QGH guides VSLNet to search for matching video span within a highlighted region.Through extensive experiments on three benchmark datasets, we show that the proposed VSLNet outperforms the state-of-the-art methods; and adopting span-based QA framework is a promising direction to solve NLVL. 1 Hao Zhang 0048, Aixin Sun, Joey Tianyi Zhou |
ACL | 1 |
| 2020 | RoSeq: Robust Sequence LabelingabstractIn this paper, we mainly investigate two issues for sequence labeling, namely, label imbalance and noisy data that are commonly seen in the scenario of named entity recognition (NER) and are largely ignored in the existing works. To address these two issues, a new method termed robust sequence labeling (RoSeq) is proposed. Specifically, to handle the label imbalance issue, we first incorporate label statistics in a novel conditional random field (CRF) loss. In addition, we design an additional loss to reduce the weights of overwhelming easy tokens for augmenting the CRF loss. To address the noisy training data, we adopt an adversarial training strategy to improve model generalization. In experiments, the proposed RoSeq achieves the state-of-the-art performances on CoNLL and English Twitter NER-88.07% on CoNLL-2002 Dutch, 87.33% on CoNLL-2002 Spanish, 52.94% on WNUT-2016 Twitter, and 43.03% on WNUT-2017 Twitter without using the additional data. Joey Tianyi Zhou, Hao Zhang 0048, Di Jin 0005, Xi Peng 0001, Yang Xiao 0007, Zhiguo Cao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | Dual Adversarial Neural Transfer for Low-Resource Named Entity RecognitionabstractWe propose a new neural transfer method termed Dual Adversarial Transfer Network (DATNet) for addressing low-resource Named Entity Recognition (NER).Specifically, two variants of DATNet, i.e., DATNet-F and DATNet-P, are investigated to explore effective feature fusion between high and low resource.To address the noisy and imbalanced training data, we propose a novel Generalized Resource-Adversarial Discriminator (GRAD).Additionally, adversarial training is adopted to boost model generalization.In experiments, we examine the effects of different components in DATNet across domains and languages, and show that significant improvement can be obtained especially for lowresource data, without augmenting any additional hand-crafted features and pre-trained language model. Joey Tianyi Zhou, Hao Zhang 0048, Di Jin 0005, Hongyuan Zhu 0002, Rick Siow Mong Goh, Kenneth Kwok |
ACL (1) | 2 |
| 2019 | Learning With Annotation of Various DegreesabstractIn this paper, we study a new problem in the scenario of sequences labeling. To be exact, we consider that the training data are with annotation of various degrees, namely, fully labeled, unlabeled, and partially labeled sequences. The learning with fully un/labeled sequence refers to the standard setting in traditional un/supervised learning, and the proposed partially labeling specifies the subject that the element does not belong to. The partially labeled data are cheaper to obtain compared with the fully labeled data though it is less informative, especially when the tasks require a lot of domain knowledge. To solve such a practical challenge, we propose a novel deep conditional random field (CRF) model which utilizes an end-to-end learning manner to smoothly handle fully/un/partially labeled sequences within a unified framework. To the best of our knowledge, this could be one of the first works to utilize the partially labeled instance for sequence labeling, and the proposed algorithm unifies the deep learning and CRF in an end-to-end framework. Extensive experiments show that our method achieves state-of-the-art performance in two sequence labeling tasks on some popular data sets. Joey Tianyi Zhou, Hao Zhang 0048, Chen Gong 0002, Xi Peng 0001, Zhiguo Cao 0001, Rick Siow Mong Goh |
IEEE Trans. Neural Networks Learn. Syst. | 3 |