EDBT 2026 Demo / reviewers in the wild / expert
Zhen-Zhong Lan
dblp:27/3780 · also Zhen-zhong Lan, Zhenzhong Lan
· DBLP profile ↗
48ranked-venue papers
7as first author
30since 2021 · last 2026
0000-0003-4763-6148ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 37 · 2 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CreditDecoding: Accelerating Parallel Decoding in Diffusion Large Language Models with Trace CreditabstractKangyu Wang, Zhiyun Jiang, Haibo Feng, Weijia Zhao, Lin Liu, Jianguo Li, Zhenzhong Lan, Weiyao Lin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Kangyu Wang, Zhiyun Jiang, Haibo Feng, Weijia Zhao, Zhen-Zhong Lan, Weiyao Lin |
ACL (1) | 7 |
| 2026 | RECAP: Resistance Capture in Text-based Mental Health Counseling with Large Language ModelsabstractRecognizing and navigating client resistance is critical for effective mental health counseling, yet detecting such behaviors is particularly challenging in text-based interactions.Existing NLP approaches oversimplify resistance categories, ignore the sequential dynamics of therapeutic interventions, and offer limited interpretability.To address these limitations, we propose Psy-FIRE, a theoretically grounded framework capturing 13 fine-grained resistance behaviors alongside collaborative interactions.Based on PsyFIRE, we construct the ClientResistance corpus with 23,930 annotated utterances from real-world Chinese text-based counseling, each supported by context-specific rationales.Leveraging this dataset, we develop RECAP, a twostage framework that detects resistance and fine-grained resistance types with explanations.RECAP achieves 91.25% F1 for distinguishing collaboration and resistance and 66.58% macro-F1 for fine-grained resistance categories classification, outperforming leading promptbased LLM baselines by over 20 points.Applied to a separate counseling dataset and a pilot study with 62 counselors, RECAP reveals the prevalence of resistance, its negative impact on therapeutic relationships and demonstrates its potential to improve counselors' understanding and intervention strategies. Anqi Li 0002, Yuqian Chen, Yi Zhu 0001, Zhen-Zhong Lan |
CoNLL | 7 |
| 2026 | Inclusion arena: A theoretically grounded framework for evaluating large foundation models via application-embedded pairwise comparisons
Hongliang He 0002, Kangyu Wang, Ruiqi Liang, Renjun Xu, Zhen-Zhong Lan |
Neurocomputing | 7 |
| 2025 | OpenWebVoyager: Building Multimodal Web Agents via Iterative Real-World Exploration, Feedback and OptimizationabstractHongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Hongming Zhang, Tianqing Fang, Zhenzhong Lan, Dong Yu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Hongliang He 0002, Wenlin Yao, Kaixin Ma, Wenhao Yu 0002, Hongming Zhang 0009, Tianqing Fang, Zhen-Zhong Lan, Dong Yu 0001 |
ACL (1) | 7 |
| 2025 | PsyDial: A Large-scale Long-term Conversational Dataset for Mental Health SupportabstractDialogue systems for mental health counseling aim to alleviate client distress and assist individuals in navigating personal challenges. Developing effective conversational agents for psychotherapy requires access to high-quality, real-world, long-term client-counselor interaction data, which is difficult to obtain due to privacy concerns. Although removing personally identifiable information is feasible, this process is labor-intensive. To address these challenges, we propose a novel privacy-preserving data reconstruction method that reconstructs real-world client-counselor dialogues while mitigating privacy concerns. We apply the RMRR (Retrieve, Mask, Reconstruct, Refine) method, which facilitates the creation of the privacy-preserving PsyDial dataset, with an average of 37.8 turns per dialogue. Extensive analysis demonstrates that PsyDial effectively reduces privacy risks while maintaining dialogue diversity and conversational exchange. To fairly and reliably evaluate the performance of models fine-tuned on our dataset, we manually collect 101 dialogues from professional counseling books. Experimental results show that models fine-tuned on PsyDial achieve improved psychological counseling performance, outperforming various baseline models. A user study involving counseling experts further reveals that our LLM-based counselor provides higher-quality responses. Code, data, and models are available at https://github.com/qiuhuachuan/PsyDial, serving as valuable resources for future advancements in AI psychotherapy. Huachuan Qiu, Zhen-Zhong Lan |
ACL (1) | 2 |
| 2025 | Value Residual LearningabstractWhile Transformer models have achieved remarkable success in various domains, the effectiveness of information propagation through deep networks remains a critical challenge.Standard hidden state residuals often fail to adequately preserve initial token-level information in deeper layers.This paper introduces ResFormer, a novel architecture that enhances information flow by incorporating value residual connections in addition to hidden state residuals.And a variant is SVFormer, where all layers share the first layer's value embedding.Comprehensive empirical evidence demonstrates ResFormer achieves equivalent validation loss with 16.11% fewer model parameters and 20.3% less training data compared to Transformer, while maintaining similar memory usage and computational cost.Besides, SVFormer reduces KV cache size by nearly half with only a small performance penalty and can be integrated with other KVefficient methods, yielding further reductions in KV cache, with performance influenced by sequence length and cumulative learning rate. Zhanchao Zhou, Zhiyun Jiang, Fares Obeid, Zhen-Zhong Lan |
ACL (1) | 5 |
| 2025 | Cognitive Representation in Large Language Models: Formalizing Psychological Constructs for Automated Questionnaire Generation
Yang Yan 0005, Lizhi Ma, Renjun Xu, Zhen-Zhong Lan |
CogSci | 5 |
| 2025 | Cognitive Distillation with Parameter-Efficient LLMs: Chain-of-Thought Calibration for Personality Prediction
Yang Yan 0005, Lizhi Ma, Renjun Xu, Zhen-Zhong Lan |
CogSci | 5 |
| 2025 | Dynamics of Instruction Fine-Tuning for Chinese Large Language ModelsabstractInstruction tuning is a burgeoning method to elicit the general intelligence of Large Language Models (LLMs). While numerous studies have examined the impact of factors such as data volume and model size on English models, the scaling properties of instruction tuning in other languages remain largely unexplored. In this work, we systematically investigate the effects of data quantity, model size, and data construction methods on instruction tuning for Chinese LLMs. We utilize a newly curated dataset, DoIT, which includes over 40,000 high-quality instruction instances covering ten underlying abilities, such as creative writing, code generation, and logical reasoning. Our experiments, conducted on models ranging from 7b to 33b parameters, yield three key findings: (i) While these factors directly affect overall model performance, some abilities are more responsive to scaling, whereas others demonstrate significant resistance. (ii) The scaling sensitivity of different abilities to these factors can be explained by two features: Complexity and Transference. (iii) By tailoring training strategies to their varying sensitivities, specific abilities can be efficiently learned, enhancing performance on two public benchmarks. Chiyu Song, Zhanchao Zhou, Jianhao Yan, Yuejiao Fei, Zhen-Zhong Lan, Yue Zhang 0004 |
COLING | 5 |
| 2025 | Do Large Language Models Truly Grasp Addition? A Rule-Focused Diagnostic Using Two-Integer ArithmeticabstractLarge language models (LLMs) achieve impressive results on advanced mathematics benchmarks but sometimes fail on basic arithmetic tasks, raising the question of whether they have truly grasped fundamental arithmetic rules or are merely relying on pattern matching.To unravel this issue, we systematically probe LLMs' understanding of two-integer addition (0 to 2 64 ) by testing three crucial properties: commutativity (A + B = B + A), representation invariance via symbolic remapping (e.g., 7 → Y), and consistent accuracy scaling with operand length.Our evaluation of 12 leading LLMs reveals a stark disconnect: while models achieve high numeric accuracy (73.8-99.8%),they systematically fail these diagnostics.Specifically, accuracy plummets to ≤ 7.5% with symbolic inputs, commutativity is violated in up to 20% of cases, and accuracy scaling is non-monotonic.Interventions further expose this pattern-matching reliance: explicitly providing rules degrades performance by 29.49%, while prompting for explanations before answering merely maintains baseline accuracy.These findings demonstrate that current LLMs address elementary addition via pattern matching, not robust rule induction, motivating new diagnostic benchmarks and innovations in model architecture and training to cultivate genuine mathematical reasoning.we release both our diagnostic dataset and the code for dataset generation at https://github.com/ kuri-leo/llm-arithmetic-diagnostic. Yang Yan 0005, Renjun Xu, Zhen-Zhong Lan |
EMNLP | 4 |
| 2025 | Enhancing Chest X-ray Classification through Knowledge Injection in Cross-Modality LearningabstractThe integration of artificial intelligence in medical imaging has shown tremendous potential, yet the relationship between pre-trained knowledge and performance in cross-modality learning remains unclear. This study investigates how explicitly injecting medical knowledge into the learning process affects the performance of cross-modality classification, focusing on Chest X-ray (CXR) images. We introduce a novel Set Theory-based knowledge injection framework that generates captions for CXR images with controllable knowledge granularity. Using this framework, we fine-tune CLIP model on captions with varying levels of medical information. We evaluate the model’s performance through zero-shot classification on the CheXpert dataset, a benchmark for CXR classification. Our results demonstrate that injecting fine-grained medical knowledge substantially improves classification accuracy, achieving 72.5% compared to 49.9% when using human-generated captions. This highlights the crucial role of domain-specific knowledge in medical cross-modality learning. Furthermore, we explore the influence of knowledge density and the use of domain-specific Large Language Models (LLMs) for caption generation, finding that denser knowledge and specialized LLMs contribute to enhanced performance. This research advances medical image analysis by demonstrating the effectiveness of knowledge injection for improving automated CXR classification, paving the way for more accurate and reliable diagnostic tools. Yang Yan 0005, Bingqing Yue, Qiaxuan Li, Man Huang, Zhen-Zhong Lan |
ICASSP | 6 |
| 2025 | ConceptPsy: A comprehensive benchmark suite for hierarchical psychological concept understanding in LLMs
Junlei Zhang, Hongliang He 0002, Lizhi Ma, Nirui Song, Shuyuan He, Huachuan Qiu, Zhanchao Zhou, Anqi Li 0002, Yong Dai 0001, Renjun Xu, Zhen-Zhong Lan |
Neurocomputing | 12 |
| 2024 | WebVoyager: Building an End-to-End Web Agent with Large Multimodal ModelsabstractHongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, Dong Yu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Hongliang He 0002, Wenlin Yao, Kaixin Ma, Wenhao Yu 0002, Yong Dai 0001, Hongming Zhang 0009, Zhen-Zhong Lan, Dong Yu 0001 |
ACL (1) | 7 |
| 2024 | PsyChat: A Client-Centric Dialogue System for Mental Health SupportabstractDialogue systems are increasingly integrated into mental health support to help clients facilitate exploration, gain insight, take action, and ultimately heal themselves. A practical and user-friendly dialogue system should be client-centric, focusing on the client’s behaviors. However, existing dialogue systems publicly available for mental health support often concentrate solely on the counselor’s strategies rather than the behaviors expressed by clients. This can lead to unreasonable or inappropriate counseling strategies and corresponding responses generated by the dialogue system. To address this issue, we propose PsyChat, a client-centric dialogue system that provides psychological support through online chat. The client-centric dialogue system comprises five modules: client behavior recognition, counselor strategy selection, input packer, response generator, and response selection. Both automatic and human evaluations demonstrate the effectiveness and practicality of our proposed dialogue system for real-life mental health support. Furthermore, the case study demonstrates that the dialogue system can predict the client’s behaviors, select appropriate counselor strategies, and generate accurate and suitable responses. Huachuan Qiu, Anqi Li 0002, Lizhi Ma, Zhen-Zhong Lan |
CSCWD | 4 |
| 2024 | Facilitating Pornographic Text Detection for Open-Domain Dialogue Systems via Knowledge Distillation of Large Language ModelsabstractPornographic content occurring in human-machine interaction dialogues can cause severe side effects for users in open-domain dialogue systems. However, research on detecting pornographic language within human-machine interaction dialogues is an important subject that is rarely studied. To advance in this direction, we introduce CENSORCHAT, a dialogue monitoring dataset aimed at detecting whether the dialogue session contains pornographic content. To this end, we collect real-life human-machine interaction dialogues in the wild and break them down into single utterances and single-turn dialogues, with the last utterance spoken by the chatbot. We propose utilizing knowledge distillation of large language models to annotate the dataset. Specifically, first, the raw dataset is annotated by four open-source large language models, with the majority vote determining the label. Second, we use ChatGPT to update the empty label from the first step. Third, to ensure the quality of the validation and test sets, we utilize GPT-4 for label calibration. If the current label does not match the one generated by GPT-4, we employ a self-criticism strategy to verify its correctness. Finally, to facilitate the detection of pornographic text, we develop a series of text classifiers using a pseudo-labeled dataset. Detailed data analysis demonstrates that leveraging knowledge distillation techniques with large language models provides a practical and cost-efficient method for developing pornographic text detectors. Huachuan Qiu, Hongliang He 0002, Anqi Li 0002, Zhen-Zhong Lan |
CSCWD | 5 |
| 2024 | Tailored Visions: Enhancing Text-to-Image Generation with Personalized Prompt RewritingabstractDespite significant progress in the field, it is still challenging to create personalized visual representations that align closely with the desires and preferences of individ-ual users. This process requires users to articulate their ideas in words that are both comprehensible to the models and accurately capture their vision, posing difficul-ties for many users. In this paper, we tackle this challenge by leveraging historical user interactions with the system to enhance user prompts. We propose a novel approach that involves rewriting user prompts based on a newly collected large-scale text-to-image dataset with over 300k prompts from 3115 users. Our rewriting model enhances the expressiveness and alignment of user prompts with their intended visual outputs. Experimental results demonstrate the superiority of our methods over baseline approaches, as evidenced in our new offline evaluation method and online tests. Our code and dataset are available at https://github.com/zzjchen/Tailored-Visions Lichao Zhang 0001, Fangsheng Weng, Lili Pan 0001, Zhen-Zhong Lan |
CVPR | 5 |
| 2024 | PsyGUARD: An Automated System for Suicide Detection and Risk Assessment in Psychological CounselingabstractAs awareness of mental health issues grows, online counseling support services are becoming increasingly prevalent worldwide.Detecting whether users express suicidal ideation in textbased counseling services is crucial for identifying and prioritizing at-risk individuals.However, the lack of domain-specific systems to facilitate fine-grained suicide detection and corresponding risk assessment in online counseling poses a significant challenge for automated crisis intervention aimed at suicide prevention.In this paper, we propose PsyGUARD, an automated system for detecting suicide ideation and assessing risk in psychological counseling.To achieve this, we first develop a detailed taxonomy for detecting suicide ideation based on foundational theories.We then curate a largescale, high-quality dataset called PsySUICIDE for suicide detection.To evaluate the capabilities of automated systems in fine-grained suicide detection, we establish a range of baselines.Subsequently, to assist automated services in providing safe, helpful, and tailored responses for further assessment, we propose to build a suite of risk assessment frameworks.Our study not only provides an insightful analysis of the effectiveness of automated risk assessment systems based on fine-grained suicide detection but also highlights their potential to improve mental health services on online counseling platforms.Code, data, and models are available at https://github.com/qiuhuachuan/ PsyGUARD. Huachuan Qiu, Lizhi Ma, Zhen-Zhong Lan |
EMNLP | 3 |
| 2024 | Benchmark for Detecting Child Pornography in Open Domain Dialogues Using Large Language ModelsabstractAs large language models become increasingly prevalent, their safe and secure application, particularly in preventing the generation of child pornographic content in real-world open-domain dialogues, has become a crucial concern. Despite the urgency of this issue, research efforts are hindered by the lack of dedicated datasets for this area. Addressing this gap, we introduce a pioneering benchmark dataset specifically designed for the detection of child pornography in open-domain dialogues. Recognizing the intrinsic complexities involved in labling such data, we developed a novel Distillation-Based Recurrent Extraction method. This approach enables us to efficiently gather, annotate, and refine the data collection process. Our dataset categorizes dialogues into three distinct sections: non-pornographic, child pornographic, and adult pornographic, ensuring clear differentiation between child and adult content. Through extensive experiments, we demonstrate that LLMs including BERT, RoBERTa, LLaMA, among others, significantly enhance their detection capabilities when fine-tuned with our dataset. This improvement not only attests to the dataset’s immediate utility but also highlights its importance for guiding future research. Furthermore, our findings indicate substantial potential for further advancements in detection performance, emphasizing the critical role of our benchmark dataset in ongoing research efforts. Zhiwei Huang 0006, Lichao Zhang 0001, Yuming Yan, Zhenyang Xiao, Zhen-Zhong Lan |
IJCNN | 7 |
| 2024 | AgentBoard: An Analytical Evaluation Board of Multi-turn LLM AgentsabstractEvaluating large language models (LLMs) as general-purpose agents is essential for understanding their capabilities and facilitating their integration into practical applications. However, the evaluation process presents substantial challenges. A primary obstacle is the benchmarking of agent performance across diverse scenarios within a unified framework, especially in maintaining partially-observable environments and ensuring multi-round interactions. Moreover, current evaluation frameworks mostly focus on the final success rate, revealing few insights during the process and failing to provide a deep understanding of the model abilities. To address these challenges, we introduce AgentBoard, a pioneering comprehensive benchmark and accompanied open-source evaluation framework tailored to analytical evaluation of LLM agents. AgentBoard offers a fine-grained progress rate metric that captures incremental advancements as well as a comprehensive evaluation toolkit that features easy assessment of agents for multi-faceted analysis through interactive visualization. This not only sheds light on the capabilities and limitations of LLM agents but also propels the interpretability of their performance to the forefront. Ultimately, AgentBoard serves as a significant step towards demystifying agent behaviors and accelerating the development of stronger LLM agents. Junlei Zhang, Cheng Yang 0007, Yujiu Yang 0001, Yaohui Jin, Zhen-Zhong Lan, Lingpeng Kong, Junxian He |
NeurIPS | 7 |
| 2024 | Offline prompt polishing for low quality instructions
Zhanchao Zhou, Yuming Yan, Renjun Xu, Zhen-Zhong Lan |
Neurocomputing | 7 |
| 2023 | Instance Smoothed Contrastive Learning for Unsupervised Sentence EmbeddingabstractContrastive learning-based methods, such as unsup-SimCSE, have achieved state-of-the-art (SOTA) performances in learning unsupervised sentence embeddings. However, in previous studies, each embedding used for contrastive learning only derived from one sentence instance, and we call these embeddings instance-level embeddings. In other words, each embedding is regarded as a unique class of its own, which may hurt the generalization performance. In this study, we propose IS-CSE (instance smoothing contrastive sentence embedding) to smooth the boundaries of embeddings in the feature space. Specifically, we retrieve embeddings from a dynamic memory buffer according to the semantic similarity to get a positive embedding group. Then embeddings in the group are aggregated by a self-attention operation to produce a smoothed instance embedding for further analysis. We evaluate our method on standard semantic text similarity (STS) tasks and achieve an average of 78.30%, 79.47%, 77.73%, and 79.42% Spearman’s correlation on the base of BERT-base, BERT-large, RoBERTa-base, and RoBERTa-large respectively, a 2.05%, 1.06%, 1.16% and 0.52% improvement compared to unsup-SimCSE. Hongliang He 0002, Junlei Zhang, Zhen-Zhong Lan, Yue Zhang 0004 |
AAAI | 3 |
| 2023 | Enhancing Grammatical Error Correction Systems with ExplanationsabstractGrammatical error correction systems improve written communication by detecting and correcting language mistakes.To help language learners better understand why the GEC system makes a certain correction, the causes of errors (evidence words) and the corresponding error types are two key factors.To enhance GEC systems with explanations, we introduce EXPECT, a large dataset annotated with evidence words and grammatical error types.We propose several baselines and analysis to understand this task.Furthermore, human evaluation verifies our explainable GEC system's explanations can assist second-language learners in determining whether to accept a correction suggestion and in understanding the associated grammar rule. Yuejiao Fei, Leyang Cui, Sen Yang 0005, Wai Lam, Zhen-Zhong Lan, Shuming Shi 0001 |
ACL (1) | 5 |
| 2023 | Understanding Client Reactions in Online Mental Health CounselingabstractAnqi Li, Lizhi Ma, Yaling Mei, Hongliang He, Shuai Zhang, Huachuan Qiu, Zhenzhong Lan. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Anqi Li 0002, Lizhi Ma, Yaling Mei, Hongliang He 0002, Huachuan Qiu, Zhen-Zhong Lan |
ACL (1) | 7 |
| 2023 | Contrastive Learning of Sentence Embeddings from ScratchabstractContrastive learning has been the dominant approach to train state-of-the-art sentence embeddings.Previous studies have typically learned sentence embeddings either through the use of human-annotated natural language inference (NLI) data or via large-scale unlabeled sentences in an unsupervised manner.However, even in the case of 1;unlabeled data, their acquisition presents challenges in certain domains due to various reasons.To address these issues, we present SynCSE, a contrastive learning framework that trains sentence embeddings with synthesized data.Specifically, we explore utilizing large language models to synthesize the required data samples for contrastive learning, including (1) producing positive and negative annotations given unlabeled sentences (SynCSE-partial), and (2) generating sentences along with their corresponding annotations from scratch (SynCSE-scratch).Experimental results on sentence similarity and reranking tasks indicate that both SynCSE-partial and SynCSE-scratch greatly outperform unsupervised baselines, and SynCSE-partial even achieves comparable performance to the supervised models in most settings.1 * Work done during Junlei's visit to HKUST.† Corresponding author. 1 Code and the synthesized datasets are available at https://github.com/hkust-nlp/SynCSE.I saw a sunset at the beach today.My city exploration led me to a beautiful building. Junlei Zhang, Zhen-Zhong Lan, Junxian He |
EMNLP | 2 |
| 2023 | A Benchmark for Understanding Dialogue Safety in Mental Health Support
Huachuan Qiu, Anqi Li 0002, Hongliang He 0002, Zhen-Zhong Lan |
NLPCC (2) | 6 |
| 2023 | Poincaré Kernels for Hyperbolic Representations
Pengfei Fang, Mehrtash Harandi, Zhen-Zhong Lan, Lars Petersson |
Int. J. Comput. Vis. | 3 |
| 2023 | Beyond a strong baseline: cross-modality contrastive learning for visible-infrared person re-identification
Pengfei Fang, Zhen-Zhong Lan |
Mach. Vis. Appl. | 3 |
| 2022 | Connecting the Dots in Self-Supervised Learning: A Brief Survey for BeginnersabstractAbstract The artificial intelligence (AI) community has recently made tremendous progress in developing self-supervised learning (SSL) algorithms that can learn high-quality data representations from massive amounts of unlabeled data. These methods brought great results even to the fields outside of AI. Due to the joint efforts of researchers in various areas, new SSL methods come out daily. However, such a sheer number of publications make it difficult for beginners to see clearly how the subject progresses. This survey bridges this gap by carefully selecting a small portion of papers that we believe are milestones or essential work. We see these researches as the “dots” of SSL and connect them through how they evolve. Hopefully, by viewing the connections of these dots, readers will have a high-level picture of the development of SSL across multiple disciplines including natural language processing, computer vision, graph learning, audio processing, and protein learning. Peng-Fei Fang, Yang Yan 0005, Qiyue Kang, Xiao-Fei Li, Zhen-Zhong Lan |
J. Comput. Sci. Technol. | 7 |
| 2021 | Do Transformer Modifications Transfer Across Implementations and Applications?abstractSharan Narang, Hyung Won Chung, Yi Tay, Liam Fedus, Thibault Fevry, Michael Matena, Karishma Malkan, Noah Fiedel, Noam Shazeer, Zhenzhong Lan, Yanqi Zhou, Wei Li, Nan Ding, Jake Marcus, Adam Roberts, Colin Raffel. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Sharan Narang, Hyung Won Chung, Yi Tay, William Fedus, Thibault Févry, Michael Matena, Karishma Malkan, Noah Fiedel, Noam Shazeer, Zhen-Zhong Lan, Yanqi Zhou, Wei Li 0133, Nan Ding 0002, Jake Marcus, Adam Roberts, Colin Raffel |
EMNLP (1) | 10 |
| 2021 | Dynamic Resolution NetworkabstractDeep convolutional neural networks (CNNs) are often of sophisticated design with numerous learnable parameters for the accuracy reason. To alleviate the expensive costs of deploying them on mobile devices, recent works have made huge efforts for excavating redundancy in pre-defined architectures. Nevertheless, the redundancy on the input resolution of modern CNNs has not been fully investigated, i.e., the resolution of input image is fixed. In this paper, we observe that the smallest resolution for accurately predicting the given image is different using the same neural network. To this end, we propose a novel dynamic-resolution network (DRNet) in which the input resolution is determined dynamically based on each input sample. Wherein, a resolution predictor with negligible computational costs is explored and optimized jointly with the desired network. Specifically, the predictor learns the smallest resolution that can retain and even exceed the original recognition accuracy for each image. During the inference, each input image will be resized to its predicted resolution for minimizing the overall computation burden. We then conduct extensive experiments on several benchmark networks and datasets. The results show that our DRNet can be embedded in any off-the-shelf network architecture to obtain a considerable reduction in computational complexity. For instance, DR-ResNet-50 achieves similar performance with an about 34% computation reduction, while gaining 1.4% accuracy increase with 10% computation reduction compared to the original ResNet-50 on ImageNet. Code will be available at https://gitee.com/mindspore/models/tree/master/research/cv/DRNet. Mingjian Zhu, Kai Han 0002, Enhua Wu, Qiulin Zhang, Zhen-Zhong Lan, Yunhe Wang 0001 |
NeurIPS | 6 |
| 2020 | CLUE: A Chinese Language Understanding Evaluation BenchmarkabstractLiang Xu, Hai Hu, Xuanwei Zhang, Lu Li, Chenjie Cao, Yudong Li, Yechen Xu, Kai Sun, Dian Yu, Cong Yu, Yin Tian, Qianqian Dong, Weitang Liu, Bo Shi, Yiming Cui, Junyi Li, Jun Zeng, Rongzhao Wang, Weijian Xie, Yanting Li, Yina Patterson, Zuoyu Tian, Yiwen Zhang, He Zhou, Shaoweihua Liu, Zhe Zhao, Qipeng Zhao, Cong Yue, Xinrui Zhang, Zhengliang Yang, Kyle Richardson, Zhenzhong Lan. Proceedings of the 28th International Conference on Computational Linguistics. 2020. Liang Xu 0011, Hai Hu 0001, Xuanwei Zhang, Chenjie Cao, Yudong Li 0001, Yechen Xu, Kai Sun 0006, Dian Yu 0001, Cong Yu 0010, Yin Tian, Qianqian Dong, Weitang Liu, Yiming Cui 0001, Rongzhao Wang, Weijian Xie, Yina Patterson, Zuoyu Tian, Shaoweihua Liu, Zhe Zhao 0006, Qipeng Zhao, Cong Yue, Zhengliang Yang, Kyle Richardson 0001, Zhen-Zhong Lan |
COLING | 32 |
| 2020 | ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
Zhen-Zhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, Radu Soricut |
ICLR | 1 |
| 2018 | Hidden Two-Stream Convolutional Networks for Action Recognition
Yi Zhu 0001, Zhen-Zhong Lan, Shawn D. Newsam, Alex Hauptmann 0001 |
ACCV (3) | 2 |
| 2015 | Beyond Gaussian Pyramid: Multi-skip Feature Stacking for action recognitionabstractMost state-of-the-art action feature extractors involve differential operators, which act as highpass filters and tend to attenuate low frequency action information. This attenuation introduces bias to the resulting features and generates ill-conditioned feature matrices. The Gaussian Pyramid has been used as a feature enhancing technique that encodes scale-invariant characteristics into the feature space in an attempt to deal with this attenuation. However, at the core of the Gaussian Pyramid is a convolutional smoothing operation, which makes it incapable of generating new features at coarse scales. In order to address this problem, we propose a novel feature enhancing technique called Multi-skIp Feature Stacking (MIFS), which stacks features extracted using a family of differential filters parameterized with multiple time skips and encodes shift-invariance into the frequency space. MIFS compensates for information lost from using differential operators by recapturing information at coarse scales. This recaptured information allows us to match actions at different speeds and ranges of motion. We prove that MIFS enhances the learnability of differential-based features exponentially. The resulting feature matrices from MIFS have much smaller conditional numbers and variances than those from conventional methods. Experimental results show significantly improved performance on challenging action recognition and event detection tasks. Specifically, our method exceeds the state-of-the-arts on Hollywood2, UCF101 and UCF50 datasets and is comparable to state-of-the-arts on HMDB51 and Olympics Sports datasets. MIFS can also be used as a speedup strategy for feature extraction with minimal or no accuracy cost. Zhen-Zhong Lan, Ming Lin 0002, Xuanchong Li, Alex Hauptmann 0001, Bhiksha Raj |
CVPR | 1 |
| 2015 | Density Corrected Sparse Recovery when R.I.P. Condition Is Broken
Ming Lin 0002, Zhen-Zhong Lan, Alex Hauptmann 0001 |
IJCAI | 2 |
| 2015 | Incremental Multimodal Query Construction for Video SearchabstractRecent improvements in content-based video search have led to systems with promising accuracy, thus opening up the possibility for interactive content-based video search to the general public. We present an interactive system based on a state-of-the-art content-based video search pipeline which enables users to do multimodal text-to-video and video-to-video search in large video collections, and to incrementally refine queries through relevance feedback and model visualization. Also, the comprehensive functionalities enhance a flexible formulation of multimodal queries with different characteristics. Quantitative and qualitative analysis shows that our system is capable of assisting users to incrementally build effective queries over complex event topics. Xiaojun Chang, Shoou-I Yu, Xingzhong Du, Xuanchong Li, Lu Jiang 0004, Zexi Mao, Zhen-Zhong Lan, Susanne Burger, Alex Hauptmann 0001 |
ICMR | 9 |
| 2015 | Learn to Recognize Actions Through Neural NetworksabstractThis research seeks to develop neural network techniques to effectively recognize actions in videos. The proposed study will lead to a deeper understanding of how neural network algorithms can help AI systems to understand motions. It will also realize an action recognition system that significantly outperforms current state-of-the-art. In addition, we will investigate the extent to which our work can be beneficial to understanding how brains perceive and analyze actions. Recent data-driven neural network approaches such as convolutional neural networks have been successful for object recognition. However, learning to recognize actions has proven to be quite a challenge due to the difficulty of getting enough labels, processing large-scale video data, and capturing motion information from videos. Therefore, we leverage effective techniques from local hand-crafted methods to help neural network algorithms learn motion features. These techniques include learning from video and optical flow volumes that follow motion trajectories, pooling features from videos played at multiple frame rates to achieve speed invariance, extending the local descriptors with normalized locations to incorporate spatial-temporal information, and a training-free re-ranking technique to exploit the relationship among classes. We also discuss a fundamental problem of whether should we learn time-aware models and what our models actually capture when we feed them temporal data. Finally, we discuss the connection of our research with how brains perceive and recognize motions. Zhen-Zhong Lan |
ACM Multimedia | 1 |
| 2014 | Viral Video Style: A Closer Look at Viral Videos on YouTubeabstractViral videos that gain popularity through the process of Internet sharing are having a profound impact on society. Existing studies on viral videos have only been on small or confidential datasets. We collect by far the largest open benchmark for viral video study called CMU Viral Video Dataset, and share it with researchers from both academia and industry. Having verified existing observations on the dataset, we discover some interesting characteristics of viral videos. Based on our analysis, in the second half of the paper, we propose a model to forecast the future peak day of viral videos. The application of our work is not only important for advertising agencies to plan advertising campaigns and estimate costs, but also for companies to be able to quickly respond to rivals in viral marketing campaigns. The proposed method is unique in that it is the first attempt to incorporate video metadata into the peak day prediction. The empirical results demonstrate that the proposed method outperforms the state-of-the-art methods, with statistically significant differences. Lu Jiang 0004, Yajie Miao, Yi Yang 0001, Zhen-Zhong Lan, Alex Hauptmann 0001 |
ICMR | 4 |
| 2014 | Resource Constrained Multimedia Event Detection
Zhen-Zhong Lan, Yi Yang 0001, Nicolas Ballas, Shoou-I Yu, Alex Hauptmann 0001 |
MMM (1) | 1 |
| 2014 | Self-Paced Learning with Diversity
Lu Jiang 0004, Deyu Meng, Shoou-I Yu, Zhen-Zhong Lan, Shiguang Shan, Alex Hauptmann 0001 |
NIPS | 4 |
| 2014 | Multimedia classification and event detection using double fusion
Zhen-Zhong Lan, Shoou-I Yu, Wei Liu 0015, Alex Hauptmann 0001 |
Multim. Tools Appl. | 1 |
| 2014 | E-LAMP: integration of innovative ideas for multimedia event detection
Yi Yang 0001, Lu Jiang 0004, Shoou-I Yu, Zhen-Zhong Lan, Zhigang Ma, Waito Sze, Ehsan Younessian, Alex Hauptmann 0001 |
Mach. Vis. Appl. | 5 |
| 2013 | Space-Time Robust Representation for Action RecognitionabstractWe address the problem of action recognition in unconstrained videos. We propose a novel content driven pooling that leverages space-time context while being robust toward global space-time transformations. Being robust to such transformations is of primary importance in unconstrained videos where the action localizations can drastically shift between frames. Our pooling identifies regions of interest using video structural cues estimated by differ ent saliency functions. To combine the different structural information, we introduce an iterative structure learning algorithm, WSVM (weighted SVM), that determines the optimal saliency layout of an action model through a sparse regularizer. A new optimization method is proposed to solve the WSVM' highly non-smooth objective function. We evaluate our approach on standard action datasets (KTH, UCF50 and HMDB). Most noticeably, the accuracy of our algorithm reaches 51.8% on the challenging HMDB dataset which outperforms the state-of-the-art of 7.3% relatively. Nicolas Ballas, Yi Yang 0001, Zhen-Zhong Lan, Bertrand Delezoide, Françoise J. Prêteux, Alex Hauptmann 0001 |
ICCV | 3 |
| 2012 | Double Fusion for Multimedia Event Detection
Zhen-Zhong Lan, Shoou-I Yu, Wei Liu 0015, Alex Hauptmann 0001 |
MMM | 1 |
| 2011 | Joint segmentation and classification of human actions in videoabstractAutomatic video segmentation and action recognition has been a long-standing problem in computer vision. Much work in the literature treats video segmentation and action recognition as two independent problems; while segmentation is often done without a temporal model of the activity, action recognition is usually performed on pre-segmented clips. In this paper we propose a novel method that avoids the limitations of the above approaches by jointly performing video segmentation and action recognition. Unlike standard approaches based on extensions of dynamic Bayesian networks, our method is based on a discriminative temporal extension of the spatial bag-of-words model that has been very popular in object recognition. The classification is performed robustly within a multi-class SVM framework whereas the inference over the segments is done efficiently with dynamic programming. Experimental results on honeybee, Weizmann, and Hollywood datasets illustrate the benefits of our approach compared to state-of-the-art methods. Minh Hoai, Zhen-Zhong Lan, Fernando De la Torre |
CVPR | 2 |
| 2009 | Extending Semi-supervised Learning Methods for Inductive Transfer LearningabstractInductive transfer learning and semi-supervised learning are two different branches of machine learning. The former tries to reuse knowledge in labeled out-of-domain instances while the later attempts to exploit the usefulness of unlabeled in-domain instances. In this paper, we bridge the two branches by pointing out that many semi-supervised learning methods can be extended for inductive transfer learning, if the step of labeling an unlabeled instance is replaced by re-weighting a diff-distribution instance. Based on this recognition, we develop a new transfer learning method, namely COITL, by extending the co-training method in semi-supervised learning. Experimental results reveal that COITL can achieve significantly higher generalization and robustness, compared with two state-of-the-art methods in inductive transfer learning. Zhen-Zhong Lan, Wei Liu 0015, Wei Bi |
ICDM | 2 |
| 2008 | Differential Evolution Based on Improved Learning Strategy
Zhen-Zhong Lan, Xiang-hu Feng |
PRICAI | 2 |
| 2008 | A novel information spread evolutionary algorithm applied to function optimizationabstractInformation spread mechanism (ISM) plays an essential role in evolutionary algorithms, forming different optimization methodologies. This paper briefly analyzes some existed ISMs and proposes a novel information spread evolutionary algorithm (NISEA). The algorithm uses a special ISM aiming at diffusing partial information of an individual to accelerate the improvement of the whole individual. Two mutation strategies are incorporated to enhance the population diversity and the selection operation is adopted to direct the evolution. Extensive experiments on 23 benchmark functions are taken to evaluate the performance of NISEA. The results are compared with those obtained by particle swarm optimization (PSO) and fast evolutionary programming (FEP), demonstrating the effectiveness and efficiency of the proposed algorithm. Zhen-Zhong Lan, Xiang-hu Feng |
SMC | 1 |