VLDB 2026 Research / reviewers in the wild / expert
Zuozhu Liu
dblp:173/9297
· DBLP profile ↗
67ranked-venue papers
4as first author
59since 2021 · last 2026
0000-0002-7816-502XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 45 · 2 first-author · 40 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 3 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Computer networks · 3 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Topology-Inspired Backward-Free Framework for Test-Time Adaptation in Medical DetectionabstractRecently, Test-Time Adaptation (TTA) has gained increasing attention in medical imaging due to its ability to improve model generalization under domain shifts without retraining. In particular, directly applying a well-trained model across various medical centers faces significant performance degradation caused by variations in equipment, operators, imaging conditions, and scanning skill levels of sonographers. Existing TTA methods either rely on parameter adaptation that increases computational cost or apply simple prediction fusion that ignores anatomical structure knowledge. To address these limitations, we propose a novel backward-free Topology-aware TTA framework named T^3 that integrates Structural Perception Modeling (SPM) and Box Regression Adaptation (BRA). SPM is implemented through an organ space heatmap generated via Gaussian kernel superposition. This heatmap encodes anatomical topology without requiring additional training or source data. BRA further improves localization and classification by fusing detection outputs based on the contribution of detected results to anatomically meaningful peak points from the heatmaps. Extensive experiments were conducted across six cross-domain scenarios, and the results demonstrate that our method achieves state-of-the-art cross-domain detection performance while maintaining high efficiency, offering a practical and robust solution for real-world medical diagnostic applications. Bin Pu, Xingguo Lv, Jiewen Yang, Lei Zhao 0013, Zuozhu Liu, Kenli Li 0001 |
AAAI | 6 |
| 2026 | CP-Router: An Uncertainty-Aware Router Between LLM and LRMabstractRecent advances in large reasoning models (LRMs) have significantly enhanced long-chain reasoning capabilities over standard large language models (LLMs). However, LRMs often produce unnecessarily lengthy outputs even for simple queries, leading to inefficiencies or even accuracy degradation compared to LLMs. To address this, we propose CP-Router, a training-free, model-agnostic routing framework that dynamically selects between an LLM and an LRM, demonstrated with multiple-choice question answering (MCQA) prompts. The routing decision is guided by the prediction uncertainty estimates derived via Conformal Prediction (CP), which provides rigorous coverage guarantees. To improve uncertainty differentiation across inputs, we introduce Full and Binary Entropy (FBE), a novel entropy-based criterion that adaptively selects the appropriate CP threshold. Experiments across MCQA and QA benchmarks—including mathematics, logical reasoning, and Chinese chemistry—demonstrate that CP-Router efficiently reduces token usage while maintaining or even improving accuracy compared to using LRM alone. We further demonstrate the generality and robustness of CP-Router by extending it to diverse model pairings beyond the LLM–LRM setting. Jiayuan Su, Fulin Lin, Zhaopeng Feng, Zhenyu Xiao, Xinlong Zhao, Zuozhu Liu, Hongwei Wang 0001 |
AAAI | 8 |
| 2026 | Beyond N-grams: A Hierarchical Reward Learning Framework for Clinically-Aware Medical Report GenerationabstractAutomatic medical report generation can greatly reduce the workload of doctors, but it is often unreliable for real-world deployment. Current methods can write formally fluent sentences but may be factually flawed, introducing serious medical errors known as clinical hallucinations, which make them untrustworthy for diagnosis. To bridge this gap, we introduce HiMed-RL, a Hierarchical Medical Reward Learning Framework designed to explicitly prioritize clinical quality. HiMed-RL moves beyond simple text matching by deconstructing reward learning into three synergistic levels: it first ensures linguistic fluency at the token-level, then enforces factual grounding at the concept-level by aligning key medical terms with expert knowledge, and finally assesses high-level diagnostic consistency at the semantic-level using a specialized LLM verifier. This hierarchical reward is implemented via a Human-inspired Dynamic Reward Adjustment, a strategy which first teaches the model to learn basic facts before progressing to more complex diagnostic reasoning. Experimentally, HiMed-3B achieves state-of-the-art performance on both in-domain and out-of-domain benchmarks, particularly on the latter, with an improvement of 10.8% over the second-best baseline. Our work provides a robust paradigm for generating reports that not only improve fluency but clinical fine-grained quality. Shujian Gao, Songtao Jiang, Haoxiang Xia, Zhaolu Kang, Yemin Wang, Zuozhu Liu |
AAAI | 9 |
| 2026 | MPA: Multimodal Prototype Augmentation for Few-Shot LearningabstractRecently, Few-shot Learning (FSL) has become a popular task that aims to recognize new classes from only a few labeled examples and has been widely applied in fields such as natural science, remote sensing, and medical images. However, most existing methods focus only on the visual modality and compute prototypes directly from raw support images, which lack comprehensive and rich multimodal information. To address these limitations, we propose a novel Multimodal Prototype Augmentation FSL framework called MPA, including LLM-based Multi-Variant Semantic Enhancement (LMSE), Hierarchical Multi-View Augmentation (HMA), and an Adaptive Uncertain Class Absorber (AUCA). LMSE leverages large language models to generate diverse paraphrased category descriptions, enriching the support set with additional semantic cues. HMA exploits both natural and multi-view augmentations to enhance feature diversity (e.g., changes in viewing distance, camera angles, and lighting conditions). AUCA models uncertainty by introducing uncertain classes via interpolation and Gaussian sampling, effectively absorbing uncertain samples. Extensive experiments on four single-domain and six cross-domain FSL benchmarks demonstrate that MPA achieves superior performance compared to existing state-of-the-art methods across most settings. Notably, MPA surpasses the second-best method by 12.29% and 24.56% in the single-domain and cross-domain setting, respectively, in the 5-way 1-shot setting. Liwen Wu, Lei Zhao 0013, Qika Lin, Shaowen Yao 0001, Zuozhu Liu, Bin Pu |
AAAI | 7 |
| 2026 | MT³: A Synergistic Multi-Task RL Framework for Specializing MLLMs in Text Image Machine TranslationabstractZhaopeng Feng, Yupu Liang, Shaosheng Cao, Jiayuan Su, Jiahan Ren, Zhijie Zhou, Wenxuan Huang, Jian Wu, Zuozhu Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhaopeng Feng, Yupu Liang, Shaosheng Cao, Jiayuan Su, Jiahan Ren, Wenxuan Huang 0001, Jian Wu 0001, Zuozhu Liu |
ACL (1) | 9 |
| 2026 | Act as you think: Reinforcing Consistent Reasoning in Medical Visual Question AnsweringabstractSongtao Jiang, Yuan Wang, Ruizhe Chen, Yan Zhang, Ruilin Luo, Bohan Lei, Yeying Jin, Sibo Song, ZhiBo Yang, Jimeng Sun, Jian Wu, Zuozhu Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Songtao Jiang, Ruizhe Chen, Yan Zhang 0004, Ruilin Luo, Bohan Lei, Yeying Jin, Sibo Song, Jimeng Sun 0001, Jian Wu 0001, Zuozhu Liu |
ACL (1) | 12 |
| 2026 | Building and Benchmarking Large Language Models for Machine Translation in Social Network Services
Hongcheng Guo, Fei Zhao 0012, Shaosheng Cao, Xinze Lyu, Zijie Meng, Yao Hu 0002, Zhoujun Li 0001, Zuozhu Liu |
ICDE | 9 |
| 2025 | KPL: Training-Free Medical Knowledge Mining of Vision-Language ModelsabstractVisual Language Models such as CLIP excel in image recognition due to extensive image-text pre-training. However, applying the CLIP inference in zero-shot classification, particularly for medical image diagnosis, faces challenges due to: 1) the inadequacy of representing image classes solely with single category names; 2) the modal gap between the visual and text spaces generated by CLIP encoders. Despite attempts to enrich disease descriptions with large language models, the lack of class-specific knowledge often leads to poor performance. In addition, empirical evidence suggests that existing proxy learning methods for zero-shot image classification on natural image datasets exhibit instability when applied to medical datasets. To tackle these challenges, we introduce the Knowledge Proxy Learning (KPL) to mine knowledge from CLIP. KPL is designed to leverage CLIP's multimodal understandings for medical image classification through Text Proxy Optimization and Multimodal Proxy Learning. Specifically, KPL retrieves image-relevant knowledge descriptions from the constructed knowledge-enhanced base to enrich semantic text proxies. It then harnesses input images and these descriptions, encoded via CLIP, to stably generate multimodal proxies that boost the zero-shot classification performance. Extensive experiments conducted on both medical and natural image datasets demonstrate that KPL enables effective zero-shot image classification, outperforming all baselines. These findings highlight the great potential in this paradigm of mining knowledge from CLIP for medical image classification and broader areas. Tianxiang Hu, Jiawei Du 0002, Ruiyuan Zhang, Joey Tianyi Zhou, Zuozhu Liu |
AAAI | 6 |
| 2025 | DiffPO: Diffusion-styled Preference Optimization for Inference Time Alignment of Large Language ModelsabstractRuizhe Chen, Wenhao Chai, Zhifei Yang, Xiaotian Zhang, Ziyang Wang, Tony Quek, Joey Tianyi Zhou, Soujanya Poria, Zuozhu Liu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Ruizhe Chen, Wenhao Chai, Zhifei Yang 0004, Tony Q. S. Quek, Joey Tianyi Zhou, Soujanya Poria, Zuozhu Liu |
ACL (1) | 9 |
| 2025 | M-MAD: Multidimensional Multi-Agent Debate for Advanced Machine Translation EvaluationabstractRecent advancements in large language models (LLMs) have given rise to the LLM-as-a-judge paradigm, showcasing their potential to deliver human-like judgments. However, in the field of machine translation (MT) evaluation, current LLM-as-a-judge methods fall short of learned automatic metrics. In this paper, we propose Multidimensional Multi-Agent Debate (M-MAD), a systematic LLM-based multi-agent framework for advanced LLM-as-a-judge MT evaluation. Our findings demonstrate that M-MAD achieves significant advancements by (1) decoupling heuristic MQM criteria into distinct evaluation dimensions for fine-grained assessments; (2) employing multi-agent debates to harness the collaborative reasoning capabilities of LLMs; (3) synthesizing dimension-specific results into a final evaluation judgment to ensure robust and reliable outcomes. Comprehensive experiments show that M-MAD not only outperforms all existing LLM-as-a-judge methods but also competes with state-of-the-art reference-based automatic metrics, even when powered by a suboptimal model like GPT-4o mini. Detailed ablations and analysis highlight the superiority of our framework design, offering a fresh perspective for LLM-as-a-judge paradigm. Our code and data are publicly available at https://github.com/SU-JIAYUAN/M-MAD. Zhaopeng Feng, Jiayuan Su, Jiamei Zheng, Jiahan Ren, Yan Zhang 0004, Jian Wu 0001, Hongwei Wang 0001, Zuozhu Liu |
ACL (1) | 8 |
| 2025 | HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language ModelsabstractSongtao Jiang, Yan Zhang, Yeying Jin, Zhihang Tang, Yangyang Wu, Yang Feng, Jian Wu, Zuozhu Liu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Songtao Jiang, Yan Zhang 0004, Yeying Jin, Zhihang Tang, Yang Feng 0011, Jian Wu 0001, Zuozhu Liu |
ACL (1) | 8 |
| 2025 | Med-GLIP: Advancing Medical Language-Image Pre-Training with Large-Scale Grounded DatasetabstractMedical image grounding aims to align natural language phrases with specific regions in medical images, serving as a foundational task for intelligent diagnosis, visual question answering (VQA), and automated report generation(MRG). However, existing research is constrained by limited modality coverage, coarse-grained annotations, and the absence of a unified, generalizable grounding framework. To address these challenges, we construct a large-scale medical grounding dataset Med-GLIP-5M comprising over 5.3 million region-level annotations across seven imaging modalities, covering diverse anatomical structures and pathological findings. The dataset supports both segmentation and grounding tasks with hierarchical region labels, ranging from organ-level boundaries to fine-grained lesions. Based on this foundation, we propose Med-GLIP, a modality-aware grounding framework trained on Med-GLIP-5M. Rather than relying on explicitly designed expert modules, Med-GLIP implicitly acquires hierarchical semantic understanding from diverse training data-enabling it to recognize multi-granularity structures, such as distinguishing lungs from pneumonia lesions. Extensive experiments demonstrate that Med-GLIP consistently outperforms state-of-the-art baselines across multiple grounding benchmarks. Furthermore, integrating its spatial outputs into downstream tasks, including medical VQA and report generation, leads to substantial performance gains. Our dataset is available at Venn2025/Med-GLIP-5M. Ziye Deng, Ruihan He, Zijie Meng, Songtao Jiang, Zuozhu Liu |
BIBM | 8 |
| 2025 | Pet-Bench: Benchmarking the Abilities of Large Language Models as E-Pets in Social Network ServicesabstractAs interest in using Large Language Models for interactive and emotionally rich experiences grows, virtual pet companionship emerges as a novel yet underexplored application. Existing approaches focus on basic pet role-playing interactions without systematically benchmarking LLMs for comprehensive companionship. In this paper, we introduce PET-BENCH, a dedicated benchmark that evaluates LLMs across both self-interaction and human-interaction dimensions. Unlike prior work, PET-BENCH emphasizes self-evolution and developmental behaviors alongside interactive engagement, offering a more realistic reflection of pet companionship. It features diverse tasks such as intelligent scheduling, memory-based dialogues, and psychological conversations, with over 7,500 interaction instances designed to simulate pet behaviors. Evaluation of 28 LLMs reveals significant performance variations linked to model size and inherent capabilities, underscoring the need for specialized optimization in this domain. PET-BENCH serves as a foundational resource for benchmarking pet-related LLM abilities and advancing emotionally immersive human-pet interactions. Hongcheng Guo, Zheyong Xie, Shaosheng Cao, Boyang Wang 0006, Weiting Liu 0001, Zheyu Ye, Zhoujun Li 0001, Zuozhu Liu, Wei Lu 0011 |
CIKM | 8 |
| 2025 | DynamicNER: A Dynamic, Multilingual, and Fine-Grained Dataset for LLM-based Named Entity RecognitionabstractHanjun Luo, Yingbin Jin, Yiran Wang, Xinfeng Li, Tong Shang, Xuecheng Liu, Ruizhe Chen, Kun Wang, Hanan Salam, Qingsong Wen, Zuozhu Liu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Hanjun Luo, Yingbin Jin, Xinfeng Li, Tong Shang, Xuecheng Liu, Ruizhe Chen, Kun Wang 0056, Hanan Salam, Qingsong Wen, Zuozhu Liu |
EMNLP | 11 |
| 2025 | Recammaster: Camera-Controlled Generative Rendering From a Single VideoabstractCamera control has been actively studied in text or image conditioned video generation tasks. However, altering camera trajectories of a given video remains under-explored, despite its importance in the field of video creation. It is non-trivial due to the extra constraints of maintaining multiple-frame appearance and dynamic synchronization. To address this, we present ReCamMaster, a camera-controlled generative video re-rendering framework that reproduces the dynamic scene of an input video at novel camera trajectories. The core innovation lies in harnessing the generative capabilities of pre-trained text-to-video models through a simple yet powerful video conditioning mechanism--its capability is often overlooked in current research. To overcome the scarcity of qualified training data, we construct a comprehensive multi-camera synchronized video dataset using Unreal Engine 5, which is carefully curated to follow real-world filming characteristics, covering diverse scenes and camera movements. It helps the model generalize to in-the-wild videos. Lastly, we further improve the robustness to diverse inputs through a meticulously designed training strategy. Extensive experiments show that our method substantially outperforms existing state-of-the-art approaches. Our method also finds promising applications in video stabilization, super-resolution, and outpainting. Our code and dataset are publicly available at: https://github.com/KwaiVGI/ReCamMaster. Jianhong Bai, Menghan Xia, Xintao Wang 0002, Lianrui Mu, Jinwen Cao, Zuozhu Liu, Haoji Hu, Xiang Bai, Pengfei Wan 0001, Di Zhang 0026 |
ICCV | 7 |
| 2025 | SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse ViewpointsabstractRecent advancements in video diffusion models demonstrate remarkable capabilities in simulating real-world dynamics and 3D consistency. This progress motivates us to explore the potential of these models to maintain dynamic consistency across diverse viewpoints, a feature highly sought after in applications like virtual filming. Unlike existing methods focused on multi-view generation of single objects for 4D reconstruction, our interest lies in generating open-world videos from arbitrary viewpoints, incorporating six degrees of freedom (6 DoF) camera poses.
To achieve this, we propose a plug-and-play module that enhances a pre-trained text-to-video model for multi-camera video generation, ensuring consistent content across different viewpoints. Specifically, we introduce a multi-view synchronization module designed to maintain appearance and geometry consistency across these viewpoints. Given the scarcity of high-quality training data, we also propose a progressive training scheme that leverages multi-camera images and monocular videos as a supplement to Unreal Engine-rendered multi-camera videos. This comprehensive approach significantly benefits our model.
Experimental results demonstrate the superiority of our proposed method over existing competitors and several baselines. Furthermore, our method enables intriguing extensions, such as re-rendering a video from multiple novel viewpoints. Project webpage: https://jianhongbai.github.io/SynCamMaster/ Jianhong Bai, Menghan Xia, Xintao Wang 0002, Ziyang Yuan, Zuozhu Liu, Haoji Hu, Pengfei Wan 0001, Di Zhang 0026 |
ICLR | 5 |
| 2025 | PAD: Personalized Alignment of LLMs at Decoding-timeabstractAligning with personalized preferences, which vary significantly across cultural, educational, and political differences, poses a significant challenge due to the computational costs and data demands of traditional alignment methods. In response, this paper presents Personalized Alignment at Decoding-time (PAD), a novel framework designed to align LLM outputs with diverse personalized preferences during the inference phase, eliminating the need for additional training. By introducing a unique personalized reward modeling strategy, this framework decouples the text generation process from personalized preferences, facilitating the generation of generalizable token-level personalized rewards. The PAD algorithm leverages these rewards to guide the decoding process, dynamically tailoring the base model’s predictions to personalized preferences. Extensive experimental results demonstrate that PAD not only outperforms existing training-based alignment methods in terms of aligning with diverse preferences but also shows significant generalizability to preferences unseen during training and scalability across different base models. This work advances the capability of LLMs to meet user needs in real-time applications, presenting a substantial step forward in personalized LLM alignment. Ruizhe Chen, Wenhao Chai, Zuozhu Liu |
ICLR | 5 |
| 2025 | FairMT-Bench: Benchmarking Fairness for Multi-turn Dialogue in Conversational LLMsabstractThe increasing deployment of large language model (LLM)-based chatbots has raised concerns regarding fairness. Fairness issues in LLMs may result in serious consequences, such as bias amplification, discrimination, and harm to minority groups. Many efforts are dedicated to evaluating and mitigating biases in LLMs. However, existing fairness benchmarks mainly focus on single-turn dialogues, while multi-turn scenarios, which better reflect real-world conversations, pose greater challenges due to conversational complexity and risk for bias accumulation. In this paper, we introduce a comprehensive benchmark for fairness of LLMs in multi-turn scenarios, **FairMT-Bench**. Specifically, We propose a task taxonomy to evaluate fairness of LLMs cross three stages: context understanding, interaction fairness, and fairness trade-offs, each comprising two tasks. To ensure coverage of diverse bias types and attributes, our multi-turn dialogue dataset FairMT-10K is constructed by integrating data from established fairness benchmarks. For evaluation, we employ GPT-4 along with bias classifiers like Llama-Guard-3, and human annotators to ensure robustness. Our experiments and analysis on FairMT-10K reveal that in multi-turn dialogue scenarios, LLMs are more prone to generating biased responses, showing significant variation in performance across different tasks and models. Based on these findings, we develop a more challenging dataset, FairMT-1K, and test 15 current state-of-the-art (SOTA) LLMs on this dataset. The results highlight the current state of fairness in LLMs and demonstrate the value of this benchmark for evaluating fairness of LLMs in more realistic multi-turn dialogue contexts. This underscores the need for future works to enhance LLM fairness and incorporate FairMT-1K in such efforts. Our code and dataset are available at https://github.com/FanZT6/FairMT-bench. Zhiting Fan, Ruizhe Chen, Tianxiang Hu, Zuozhu Liu |
ICLR | 4 |
| 2025 | Modality-Fair Preference Optimization for Trustworthy MLLM AlignmentabstractMultimodal large language models (MLLMs) have achieved remarkable success across various tasks. However, separate training of visual and textual encoders often results in a misalignment of the modality. Such misalignment may lead models to generate content that is absent from the input image, a phenomenon referred to as hallucination. These inaccuracies severely undermine the trustworthiness of MLLMs in real-world applications. Despite attempts to optimize text preferences to mitigate this issue, our initial investigation indicates that the trustworthiness of MLLMs remains inadequate. Specifically, these models tend to provide preferred answers even when the input image is heavily distorted. Analysis of visual token attention also indicates that the model focuses primarily on the surrounding context rather than the key object referenced in the question. These findings highlight a misalignment between the modalities, where answers inadequately leverage input images. Motivated by our findings, we propose Modality-Fair Preference Optimization (MFPO), which comprises three components: the construction of a multimodal preference dataset in which dispreferred images differ from originals solely in key regions; an image reward loss function encouraging the model to generate answers better aligned with the input images; and an easy-to-hard iterative alignment strategy to stabilize joint modality training. Extensive experiments on three trustworthiness benchmarks demonstrate that MFPO significantly enhances the trustworthiness of MLLMs. In particular, it enables the 7B models to attain trustworthiness levels on par with, or even surpass, those of the 13B, 34B, and larger models. Songtao Jiang, Yan Zhang 0004, Ruizhe Chen, Tianxiang Hu, Yeying Jin, Qinglin He, Yang Feng 0011, Jian Wu 0001, Zuozhu Liu |
IJCAI | 9 |
| 2025 | R-LLaVA: Improving Med-VQA Understanding through Visual Region of InterestabstractArtificial intelligence has made significant strides in medical visual question answering (Med-VQA), yet prevalent studies often interpret images holistically, overlooking the visual regions of interest that may contain crucial information, potentially aligning with a doctor’s prior knowledge that can be incorporated with minimal annotations (e.g., bounding boxes). To address this gap, this paper introduces R-LLaVA, designed to enhance biomedical VQA understanding by integrating simple medical annotations as prior knowledge directly into the image space through CLIP. These annotated visual regions of interest are then fed into the LLaVA model during training, aiming to enrich the model’s understanding of biomedical queries. Experimental evaluation on four standard Med-VQA datasets demonstrates R-LLaVA’s superiority over existing state-of-the-art (SoTA) methods. Additionally, to verify the model’s capability in visual comprehension, a novel multiple-choice medical visual understanding dataset is introduced, confirming the positive impact of focusing on visual regions of interest in advancing biomedical VQA understanding. Xupeng Chen, Zhixin Lai, Kangrui Ruan, Shichu Chen, Zuozhu Liu |
IJCNN | 6 |
| 2025 | Knowing or Guessing? Robust Medical Visual Question Answering via Joint Consistency and Contrastive Learning
Songtao Jiang, Sibo Song, Yan Zhang 0004, Yeying Jin, Yang Feng 0011, Jian Wu 0001, Zuozhu Liu |
MICCAI (11) | 8 |
| 2025 | V2T-CoT: From Vision to Text Chain-of-Thought for Medical Reasoning and Diagnosis
Shujian Gao, Zhihang Tang, Xiaotang Gai, Jian Wu 0001, Zuozhu Liu |
MICCAI (5) | 8 |
| 2025 | Fair-MoE: Medical Fairness-Oriented Mixture of Experts in Vision-Language Models
Peiran Wang, Linjie Tong, Jian Wu 0001, Zuozhu Liu |
MICCAI (5) | 5 |
| 2025 | UniEdit: A Unified Tuning-Free Framework for Video Motion and Appearance EditingabstractRecent advances in text-guided video editing have showcased promising results in appearance editing (e.g., stylization). However, video motion editing in the temporal dimension (e.g., from eating to waving), which distinguishes video editing from image editing, is underexplored. In this work, we present UniEdit, a tuning-free framework that supports both video motion and appearance editing by harnessing the power of a pre-trained text-to-video generator within an inversion-then-generation framework. To realize motion editing while preserving source video content, based on the insights that temporal and spatial self-attention layers encode inter-frame and intra-frame dependency, we introduce auxiliary motion-reference and reconstruction branches to produce text-guided motion and source features respectively. The obtained features are then injected into the main editing path via temporal and spatial self-attention layers. We also validate the effectiveness and flexibility of UniEdit by deploying it on three T2V generative models with different architectures. Experiments demonstrate that UniEdit covers video motion editing and various appearance editing scenarios, and surpasses the state-of-the-art methods. Our code is publicly available. Jianhong Bai, Tianyu He, Yuchi Wang, Junliang Guo, Haoji Hu, Zuozhu Liu, Jiang Bian 0002 |
ACM Multimedia | 6 |
| 2025 | 3D-RAD: A Comprehensive 3D Radiology Med-VQA Dataset with Multi-Temporal Analysis and Diverse Diagnostic TasksabstractMedical Visual Question Answering (Med-VQA) holds significant potential for clinical decision support, yet existing efforts primarily focus on 2D imaging with limited task diversity. This paper presents 3D-RAD, a large-scale dataset designed to advance 3D Med-VQA using radiology CT scans. The 3D-RAD dataset encompasses six diverse VQA tasks: anomaly detection, image observation, medical computation, existence detection, static temporal diagnosis, and longitudinal temporal diagnosis. It supports both open- and closed-ended questions while introducing complex reasoning challenges, including computational tasks and multi-stage temporal analysis, to enable comprehensive benchmarking. Extensive evaluations demonstrate that existing vision-language models (VLMs), especially medical VLMs exhibit limited generalization, particularly in multi-temporal tasks, underscoring the challenges of real-world 3D diagnostic reasoning. To drive future advancements, we release a high-quality training set 3D-RAD-T of 136,195 expert-aligned samples, showing that fine-tuning on this dataset could significantly enhance model performance. Our dataset and code, aiming to catalyze multimodal medical AI research and establish a robust foundation for 3D medical visual understanding, are publicly available. Xiaotang Gai, Zijie Meng, Jian Wu 0001, Zuozhu Liu |
NeurIPS | 6 |
| 2025 | Beyond Modality Collapse: Representation Blending for Multimodal Dataset DistillationabstractMultimodal Dataset Distillation (MDD) seeks to condense large-scale image-text datasets into compact surrogates while retaining their effectiveness for cross-modal learning. Despite recent progress, existing MDD approaches often suffer from ***Modality Collapse***, characterized by over-concentrated intra-modal representations and enlarged distributional gap across modalities. In this paper, at the first time, we identify this issue as stemming from a fundamental conflict between the over-compression behavior inherent in dataset distillation and the cross-modal supervision imposed by contrastive objectives. To alleviate modality collapse, we introduce **RepBlend**, a novel MDD framework that weakens overdominant cross-modal supervision via representation blending, thereby significantly enhancing intra-modal diversity. Additionally, we observe that current MDD methods impose asymmetric supervision across modalities, resulting in biased optimization. To address this, we propose symmetric projection trajectory matching, which synchronizes the optimization dynamics using modality-specific projection heads, thereby promoting balanced supervision and enhancing cross-modal alignment.
Experiments on Flickr-30K and MS-COCO show that RepBlend consistently outperforms prior state-of-the-art MDD methods, achieving significant gains in retrieval performance (e.g., +9.4 IR@10, +6.3 TR@10 under the 100-pair setting) and offering up to 6.7$\times$ distillation speedup. Xin Zhang 0092, Ziruo Zhang, Jiawei Du 0002, Zuozhu Liu, Joey Tianyi Zhou |
NeurIPS | 4 |
| 2025 | PaRT: Enhancing Proactive Social Chatbots with Personalized Real-Time RetrievalabstractSocial chatbots have become essential companions in daily scenarios ranging from emotional support to personal interaction. However, conventional chatbots with passive response mechanisms usually rely on users to initiate or sustain dialogues by bringing up new topics, resulting in diminished engagement and shortened dialogue duration. In this paper, we present PaRT, a novel framework enabling context-aware proactive dialogues for social chatbots through personalized real-time retrieval and generation. Specifically, PaRT first integrates user profiles and dialogue context into a large language model (LLM), which is initially prompted to refine user queries and recognize underlying intents for the upcoming conversation. Guided by refined intents, the LLM generates personalized dialogue topics as targeted queries to retrieve relevant passages from RedNote. Finally, we prompt LLMs with summarized passages to generate knowledge-grounded and engagement-optimized responses. Our approach has been running stably in a real-world production environment for more than 30 days, achieving a 21.77% improvement in the average duration of dialogues. Zihan Niu, Zheyong Xie, Shaosheng Cao, Chonggang Lu, Zheyu Ye, Tong Xu 0001, Zuozhu Liu, Yan Gao 0017, Jia Chen 0003, Yao Hu 0002 |
SIGIR | 7 |
| 2025 | Towards normalized clinical information extraction in Chinese radiology report with large language models
Qinwei Xu, Xingkun Xu, Chenyi Zhou, Zuozhu Liu, Feiyue Huang, Shaoxin Li 0004, Lifeng Zhu, Zhian Bai, Yuchen Xu 0008, Weiguo Hu |
Expert Syst. Appl. | 4 |
| 2025 | Cross-center Model Adaptive Tooth segmentation
Ruizhe Chen, Jianfei Yang 0001, Huimin Xiong, Ruiling Xu, Yang Feng 0011, Jian Wu 0001, Zuozhu Liu |
Medical Image Anal. | 7 |
| 2025 | Robust Privacy-Preserving Recommendation Systems Driven by Multimodal Federated LearningabstractRecommendation system (RS) is an important information filtering tool in nowadays digital era. With the growing concern on privacy, deploying RSs in a federated learning (FL) manner emerges as a promising solution, which can train a high-quality model on the premise that the server does not directly access sensitive user data. Nevertheless, some malicious clients can deduce user data by analyzing the uploaded model parameters. Even worse, some Byzantine clients can also send contaminated data to the server, causing blockage or failure of model convergence. In addition, most existing researches on federated recommendation algorithms only focus on unimodality learning, ignoring the assistance of multiple modality data to promote recommendation accuracy. Therefore, this article designs an FL-based privacy-preserving multimodal RS framework. To distinguish various modality data, an attention mechanism is introduced, wherein different weight ratios are assigned to various modal features. To further strengthen the privacy, local differential privacy (LDP) and personalized FL strategies are designed to identify malicious clients and bolster the resilience against Byzantine attacks. Finally, two multimodal datasets are established to verify the effectiveness of the proposed algorithm. The superiority of our proposed techniques is confirmed by the simulation results. Chenyuan Feng, Daquan Feng, Guanxin Huang, Zuozhu Liu, Zhenzhong Wang, Xiang-Gen Xia 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | LETA: Tooth Alignment Prediction Based on Dual-branch Latent EncodingabstractAccurately determining the clinical positions for each tooth is essential in orthodontics, while most existing solutions heavily rely on inefficient manual design. In this paper, we present the LETA, a dual-branch Latent Encoding based 3D Tooth Alignment. Our system takes as input the segmented individual 3D tooth meshes in the Intra-oral Scanner (IOS) dental surfaces, and automatically predicts the proper 3D pose transformation for each tooth. LETA includes three components: an Encoder that learns a latent code of dental pointcloud, a Projector that transforms the latent code of misaligned teeth to predicted aligned ones, and a Solver to estimate the transformation between different dental latent codes. A key novelty of LETA is that we extract the features from the ground truth (GT) aligned teeth to guide network learning during training. To effectively learn tooth features, our Encoder employs an improved point-wise convolutional operation and an attention-based network to extract local shape features and global context features respectively. Extensive experimental results on a large-scale dataset with 9,868 IOS surfaces demonstrate that LETA can achieve state-of-the-art performance. A further clinical applicability study reveals that our method can reduce orthodontists' workload over 60% compared to starting tooth alignment from scratch, demonstrating the strong potential of deep learning for future digital dentistry. Zefeng Shi, Zijie Meng, Ruizhe Chen, Yang Feng 0011, Jin Hao, Bing Fang, Zuozhu Liu, Youyi Zheng |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2024 | Robustness-Guided Image Synthesis for Data-Free QuantizationabstractQuantization has emerged as a promising direction for model compression. Recently, data-free quantization has been widely studied as a promising method to avoid privacy concerns, which synthesizes images as an alternative to real training data. Existing methods use classification loss to ensure the reliability of the synthesized images. Unfortunately, even if these images are well-classified by the pre-trained model, they still suffer from low semantics and homogenization issues. Intuitively, these low-semantic images are sensitive to perturbations, and the pre-trained model tends to have inconsistent output when the generator synthesizes an image with low semantics. To this end, we propose Robustness-Guided Image Synthesis (RIS), a simple but effective method to enrich the semantics of synthetic images and improve image diversity, further boosting the performance of data-free compression tasks. Concretely, we first introduce perturbations on input and model weight, then define the inconsistency metrics at feature and prediction levels before and after perturbations. On the basis of inconsistency on two levels, we design a robustness optimization objective to eliminate low-semantic images. Moreover, we also make our approach diversity-aware by forcing the generator to synthesize images with small correlations. With RIS, we achieve state-of-the-art performance for various settings on data-free quantization and can be extended to other data-free compression tasks. Jianhong Bai, Huanpeng Chu, Hualiang Wang, Zuozhu Liu, Ruizhe Chen, Xiaoxuan He, Lianrui Mu, Chengfei Cai, Haoji Hu |
AAAI | 5 |
| 2024 | BiasAlert: A Plug-and-play Tool for Social Bias Detection in LLMsabstractEvaluating the bias in Large Language Models (LLMs) becomes increasingly crucial with their rapid development.However, existing evaluation methods rely on fixed-form outputs and cannot adapt to the flexible open-text generation scenarios of LLMs (e.g., sentence completion and question answering).To address this, we introduce BiasAlert, a plug-and-play tool designed to detect social bias in open-text generations of LLMs.BiasAlert integrates external human knowledge with inherent reasoning capabilities to detect bias reliably.Extensive experiments demonstrate that BiasAlert significantly outperforms existing state-of-theart methods like GPT4-as-A-Judge in detecting bias.Furthermore, through application studies, we demonstrate the utility of BiasAlert in reliable LLM bias evaluation and bias mitigation across various scenarios.Model and code will be publicly released. Zhiting Fan, Ruizhe Chen, Ruiling Xu, Zuozhu Liu |
EMNLP | 4 |
| 2024 | Ladder: A Model-Agnostic Framework Boosting LLM-based Machine Translation to the Next LevelabstractGeneral-purpose Large Language Models (LLMs) like GPT-4 have achieved remarkable advancements in machine translation (MT) by leveraging extensive web content.On the other hand, translation-specific LLMs are built by pre-training on domain-specific monolingual corpora and fine-tuning with human-annotated translation data.Despite the superior performance, these methods either demand an unprecedented scale of computing and data or substantial human editing and annotation efforts.In this paper, we develop MT-Ladder, a novel model-agnostic and cost-effective tool to refine the performance of general LLMs for MT.MT-Ladder is trained on pseudo-refinement triplets which can be easily obtained from existing LLMs without additional human cost.During training, we propose a hierarchical finetuning strategy with an easy-to-hard schema, improving MT-Ladder's refining performance progressively.The trained MT-Ladder can be seamlessly integrated with any general-purpose LLMs to boost their translation performance.By utilizing Gemma-2B/7B as the backbone, MT-Ladder-2B can elevate raw translations to the level of top-tier open-source models (e.g., refining BigTranslate-13B with +6.91 BLEU and +3.52 COMET for XX→En), and MT-Ladder-7B can further enhance model performance to be on par with the state-of-theart GPT-4.Extensive ablation and analysis corroborate the effectiveness of MT-Ladder in diverse settings.Our code is available at https://github.com/fzp0424/MT-Ladder. Zhaopeng Feng, Ruizhe Chen, Yan Zhang 0004, Zijie Meng, Zuozhu Liu |
EMNLP | 5 |
| 2024 | MedCoT: Medical Chain of Thought via Hierarchical ExpertabstractArtificial intelligence has advanced in Medical Visual Question Answering (Med-VQA), but prevalent research tends to focus on the accuracy of the answers, often overlooking the reasoning paths and interpretability, which are crucial in clinical settings.Besides, current Med-VQA algorithms, typically reliant on singular models, lack the robustness needed for real-world medical diagnostics which usually require collaborative expert evaluation.To address these shortcomings, this paper presents MedCoT, a novel hierarchical expert verification reasoning chain method designed to enhance interpretability and accuracy in biomedical imaging inquiries.MedCoT is predicated on two principles: The necessity for explicit reasoning paths in Med-VQA and the requirement for multi-expert review to formulate accurate conclusions.The methodology involves an Initial Specialist proposing diagnostic rationales, followed by a Follow-up Specialist who validates these rationales, and finally, a consensus is reached through a vote among a sparse Mixture of Experts within the locally deployed Diagnostic Specialist, which then provides the definitive diagnosis.Experimental evaluations on four standard Med-VQA datasets demonstrate that MedCoT surpasses existing state-of-the-art approaches, providing significant improvements in performance and interpretability.Code is released at https: //github.com/JXLiu-AI/MedCoT. Jiawei Du 0002, Joey Tianyi Zhou, Zuozhu Liu |
EMNLP | 5 |
| 2024 | DynaThink: Fast or Slow? A Dynamic Decision-Making Framework for Large Language ModelsabstractLarge language models (LLMs) have demonstrated emergent capabilities across diverse reasoning tasks via popular Chains-of-Thought (COT) prompting.However, such a simple and fast COT approach often encounters limitations in dealing with complicated problems, while a thorough method, which considers multiple reasoning pathways and verifies each step carefully, results in slower inference.This paper addresses the challenge of enabling LLMs to autonomously select between fast and slow inference methods, thereby optimizing both efficiency and effectiveness.We introduce a dynamic decision-making framework that categorizes tasks into two distinct pathways: 'Fast', designated for tasks where the LLM quickly identifies a high-confidence solution, and 'Slow', allocated for tasks that the LLM perceives as complex and for which it has low confidence in immediate solutions as well as requiring more reasoning paths to verify.Experiments on five popular reasoning benchmarks demonstrated the superiority of the Dyna-Think over baselines. Yan Zhang 0004, Chen Zhang 0020, Zuozhu Liu, Hongwei Wang 0001, Haizhou Li 0001 |
EMNLP | 4 |
| 2024 | Multimodal Survival Ensemble Network: Integrating Genomic and Histopathological Insights for Enhanced Cancer PrognosisabstractCancer’s inherent heterogeneity demands a multimodal approach to provide an accurate prognosis, taking into account histological, clinical, and genomic data. As the field of artificial intelligence evolves with advancements in multimodal learning, its role in survival analysis becomes increasingly critical. We introduce the Multimodal Survival Ensemble Network (MSEN), a novel weakly-supervised framework designed for the seamless integration of genomic data and histopathological images. Not only does our method preserve the heterogeneity among different genomic modalities during integration, but it also ensures superior retention of spatial information in histopathological images compared to traditional techniques. Rigorous evaluations across five datasets highlight MSEN’s superior performance, marking a progressive step in cancer prognosis. Chenyi Zhou, Hualiang Wang, Xiaomeng Li 0001, Wanlu Liu, Zuozhu Liu |
ICASSP | 5 |
| 2024 | FedLoGe: Joint Local and Generic Federated Learning under Long-tailed DataabstractFederated Long-Tailed Learning (Fed-LT), a paradigm wherein data collected from decentralized local clients manifests a globally prevalent long-tailed distribution, has garnered considerable attention in recent times. In the context of Fed-LT, existing works have predominantly centered on addressing the data imbalance issue to enhance the efficacy of the generic global model while neglecting the performance at the local level. In contrast, conventional Personalized Federated Learning (pFL) techniques are primarily devised to optimize personalized local models under the presumption of a balanced global data distribution. This paper introduces an approach termed Federated Local and Generic Model Training in Fed-LT (FedLoGe), which enhances both local and generic model performance through the integration of representation learning and classifier alignment within a neural collapse framework. Our investigation reveals the feasibility of employing a shared backbone as a foundational framework for capturing overarching global trends, while concurrently employing individualized classifiers to encapsulate distinct refinements stemming from each client’s local features. Building upon this discovery, we establish the Static Sparse Equiangular Tight Frame Classifier (SSE-C), inspired by neural collapse principles that naturally prune extraneous noisy features and foster the acquisition of potent data representations. Furthermore, leveraging insights from imbalance neural collapse's classifier norm patterns, we develop Global and Local Adaptive Feature Realignment (GLA-FR) via an auxiliary global classifier and personalized Euclidean norm transfer to align global features with client preferences. Extensive experimental results on CIFAR-10/100-LT, ImageNet, and iNaturalist demonstrate the advantage of our method over state-of-the-art pFL and Fed-LT approaches. Zikai Xiao, Zihan Chen 0001, Liyinglan Liu, Yang Feng 0011, Joey Tianyi Zhou, Jian Wu 0001, Wanlu Liu, Howard H. Yang, Zuozhu Liu |
ICLR | 9 |
| 2024 | PX2Tooth: Reconstructing the 3D Point Cloud Teeth from a Single Panoramic X-Ray
Huikai Wu, Zikai Xiao, Yang Feng 0011, Jian Wu 0001, Zuozhu Liu |
MICCAI (3) | 6 |
| 2024 | DIVOTrack: A Novel Dataset and Baseline Method for Cross-View Multi-Object Tracking in DIVerse Open Scenes
Shengyu Hao, Peiyuan Liu, Yibing Zhan, Kaixun Jin, Zuozhu Liu, Mingli Song, Jenq-Neng Hwang, Gaoang Wang |
Int. J. Comput. Vis. | 5 |
| 2024 | Knowledge-guided pre-training and fine-tuning: Video representation learning for action recognition
Guanhong Wang, Zhanhao He, Keyu Lu, Yang Feng 0011, Zuozhu Liu, Gaoang Wang |
Neurocomputing | 6 |
| 2024 | Accurate estimation of 6-DoF tooth pose in 3D intraoral scans for dental applications using deep learningabstractA critical step in digital dentistry is to accurately and automatically characterize the orientation and position of individual teeth, which can subsequently be used for treatment planning and simulation in orthodontic tooth alignment. This problem remains challenging because the geometric features of different teeth are complicated and vary significantly, while a reliable large-scale dataset is yet to be constructed. In this paper we propose a novel method for automatic tooth orientation estimation by formulating it as a six-degree-of-freedom (6-DoF) tooth pose estimation task. Regarding each tooth as a three-dimensional (3D) point cloud, we design a deep neural network with a feature extractor backbone and a two-branch estimation head for tooth pose estimation. Our model, trained with a novel loss function on the newly collected large-scale dataset (10 393 patients with 280 611 intraoral tooth scans), achieves an average Euler angle error of only 4.780°–5.979° and a translation L1 error of 0.663 mm on a hold-out set of 2598 patients (77 870 teeth). Comprehensive experiments show that 98.29% of the estimations produce a mean angle error of less than 15°, which is acceptable for many clinical and industrial applications. Wanghui Ding, Mengfei Yu, Hangzheng Lin, Yang Feng 0011, Zuozhu Liu |
Frontiers Inf. Technol. Electron. Eng. | 7 |
| 2024 | A Transformer-Based Knowledge Distillation Network for Cortical Cataract GradingabstractCortical cataract, a common type of cataract, is particularly difficult to be diagnosed automatically due to the complex features of the lesions. Recently, many methods based on edge detection or deep learning were proposed for automatic cataract grading. However, these methods suffer a large performance drop in cortical cataract grading due to the more complex cortical opacities and uncertain data. In this paper, we propose a novel Transformer-based Knowledge Distillation Network, called TKD-Net, for cortical cataract grading. To tackle the complex opacity problem, we first devise a zone decomposition strategy to extract more refined features and introduce special sub-scores to consider critical factors of clinical cortical opacity assessment (location, area, density) for comprehensive quantification. Next, we develop a multi-modal mix-attention Transformer to efficiently fuse sub-scores and image modality for complex feature learning. However, obtaining the sub-score modality is a challenge in the clinic, which could cause the modality missing problem instead. To simultaneously alleviate the issues of modality missing and uncertain data, we further design a Transformer-based knowledge distillation method, which uses a teacher model with perfect data to guide a student model with modality-missing and uncertain data. We conduct extensive experiments on a dataset of commonly-used slit-lamp images annotated by the LOCS III grading system to demonstrate that our TKD-Net outperforms state-of-the-art methods, as well as the effectiveness of its key components. Codes are available at https://github.com/wjh892521292/Cataract_TKD-Net. Haochao Ying, Tingting Chen 0002, Zuozhu Liu, Danny Ziyi Chen, Ke Yao, Jian Wu 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2023 | Empirical Study of Zero-Shot NER with ChatGPTabstractLarge language models (LLMs) exhibited powerful capability in various natural language processing tasks.This work focuses on exploring LLM performance on zero-shot information extraction, with a focus on the ChatGPT and named entity recognition (NER) task.Inspired by the remarkable reasoning capability of LLM on symbolic and arithmetic reasoning, we adapt the prevalent reasoning methods to NER and propose reasoning strategies tailored for NER.First, we explore a decomposed question-answering paradigm by breaking down the NER task into simpler subproblems by labels.Second, we propose syntactic augmentation to stimulate the model's intermediate thinking in two ways: syntactic prompting, which encourages the model to analyze the syntactic structure itself, and tool augmentation, which provides the model with the syntactic information generated by a parsing tool.Besides, we adapt self-consistency to NER by proposing a two-stage majority voting strategy, which first votes for the most consistent mentions, then the most consistent types.The proposed methods achieve remarkable improvements for zero-shot NER across seven benchmarks, including Chinese and English datasets, and on both domainspecific and general-domain scenarios.In addition, we present a comprehensive analysis of the error types with suggestions for optimization directions.We also verify the effectiveness of the proposed methods on the few-shot setting and other LLMs. 1 * Corresponding authors. 1 Code available at: https://github.com/Emma1066/ Zero-Shot-NER-with-ChatGPT Input Text: The player who temporarily ranks second is German athlete Bao Lizzo, with a total score of 355.02 points, slightly lower than Lanwei.Gold Label: {"German":"Geo-Political Entity", "Lanwei": "Person", "BaoꞏLizzo": "Person"} Vanilla Ans: {"German athlete Bao Lizzo": "Person", "Lanwei": "Person"} TS-SC Ans: {"BaoꞏLizzo": "Person": "Person", "Lanwei": "Person", "German": "Geo-Political Entity"} ---------------------------------- Tingyu Xie, Qi Li 0042, Jian Zhang 0083, Yan Zhang 0004, Zuozhu Liu, Hongwei Wang 0001 |
EMNLP | 5 |
| 2023 | On the Effectiveness of Out-of-Distribution Data in Self-Supervised Long-Tail Learning
Jianhong Bai, Zuozhu Liu, Hualiang Wang, Jin Hao, Yang Feng 0011, Huanpeng Chu, Haoji Hu |
ICLR | 2 |
| 2023 | TSegFormer: 3D Tooth Segmentation in Intraoral Scans with Geometry Guided Transformer
Huimin Xiong, Kunle Li, Kaiyuan Tan, Yang Feng 0011, Joey Tianyi Zhou, Jin Hao, Haochao Ying, Jian Wu 0001, Zuozhu Liu |
MICCAI (6) | 9 |
| 2023 | Towards Distribution-Agnostic Generalized Category DiscoveryabstractData imbalance and open-ended distribution are two intrinsic characteristics of the real visual world. Though encouraging progress has been made in tackling each challenge separately, few works dedicated to combining them towards real-world scenarios. While several previous works have focused on classifying close-set samples and detecting open-set samples during testing, it's still essential to be able to classify unknown subjects as human beings. In this paper, we formally define a more realistic task as distribution-agnostic generalized category discovery (DA-GCD): generating fine-grained predictions for both close- and open-set classes in a long-tailed open-world setting. To tackle the challenging problem, we propose a Self-**Ba**lanced **Co**-Advice co**n**trastive framework (BaCon), which consists of a contrastive-learning branch and a pseudo-labeling branch, working collaboratively to provide interactive supervision to resolve the DA-GCD task. In particular, the contrastive-learning branch provides reliable distribution estimation to regularize the predictions of the pseudo-labeling branch, which in turn guides contrastive learning through self-balanced knowledge transfer and a proposed novel contrastive loss. We compare BaCon with state-of-the-art methods from two closely related fields: imbalanced semi-supervised learning and generalized category discovery. The effectiveness of BaCon is demonstrated with superior performance over all baselines and comprehensive analysis across various datasets. Our code is publicly available. Jianhong Bai, Zuozhu Liu, Hualiang Wang, Ruizhe Chen, Lianrui Mu, Xiaomeng Li 0001, Joey Tianyi Zhou, Yang Feng 0011, Jian Wu 0001, Haoji Hu |
NeurIPS | 2 |
| 2023 | Fast Model DeBias with Machine UnlearningabstractRecent discoveries have revealed that deep neural networks might behave in a biased manner in many real-world scenarios. For instance, deep networks trained on a large-scale face recognition dataset CelebA tend to predict blonde hair for females and black hair for males. Such biases not only jeopardize the robustness of models but also perpetuate and amplify social biases, which is especially concerning for automated decision-making processes in healthcare, recruitment, etc., as they could exacerbate unfair economic and social inequalities among different groups. Existing debiasing methods suffer from high costs in bias labeling or model re-training, while also exhibiting a deficiency in terms of elucidating the origins of biases within the model. To this respect, we propose a fast model debiasing method (FMD) which offers an efficient approach to identify, evaluate and remove biases inherent in trained models. The FMD identifies biased attributes through an explicit counterfactual concept and quantifies the influence of data samples with influence functions. Moreover, we design a machine unlearning-based strategy to efficiently and effectively remove the bias in a trained model with a small counterfactual dataset.
Experiments on the Colored MNIST, CelebA, and Adult Income datasets demonstrate that our method achieves superior or competing classification accuracies compared with state-of-the-art retraining-based methods while attaining significantly fewer biases and requiring much less debiasing cost. Notably, our method requires only a small external dataset and updating a minimal amount of model parameters, without the requirement of access to training data that may be too large or unavailable in practice. Ruizhe Chen, Huimin Xiong, Jianhong Bai, Tianxiang Hu, Jin Hao, Yang Feng 0011, Joey Tianyi Zhou, Jian Wu 0001, Zuozhu Liu |
NeurIPS | 10 |
| 2023 | Fed-GraB: Federated Long-tailed Learning with Self-Adjusting Gradient BalancerabstractData privacy and long-tailed distribution are the norms rather than the exception in many real-world tasks. This paper investigates a federated long-tailed learning (Fed-LT) task in which each client holds a locally heterogeneous dataset; if the datasets can be globally aggregated, they jointly exhibit a long-tailed distribution. Under such a setting, existing federated optimization and/or centralized long-tailed learning methods hardly apply due to challenges in (a) characterizing the global long-tailed distribution under privacy constraints and (b) adjusting the local learning strategy to cope with the head-tail imbalance. In response, we propose a method termed $\texttt{Fed-GraB}$, comprised of a Self-adjusting Gradient Balancer (SGB) module that re-weights clients' gradients in a closed-loop manner, based on the feedback of global long-tailed distribution evaluated by a Direct Prior Analyzer (DPA) module. Using $\texttt{Fed-GraB}$, clients can effectively alleviate the distribution drift caused by data heterogeneity during the model training process and obtain a global model with better performance on the minority classes while maintaining the performance of the majority classes. Extensive experiments demonstrate that $\texttt{Fed-GraB}$ achieves state-of-the-art performance on representative datasets such as CIFAR-10-LT, CIFAR-100-LT, ImageNet-LT, and iNaturalist. Zikai Xiao, Zihan Chen 0001, Songshang Liu, Hualiang Wang, Yang Feng 0011, Jin Hao, Joey Tianyi Zhou, Jian Wu 0001, Howard H. Yang, Zuozhu Liu |
NeurIPS | 10 |
| 2023 | A Fine-Grained Attention Model for High Accuracy Operational Robot GuidanceabstractDeep learning enhanced Internet of Things (IoT) is advancing the transformation toward smart manufacturing. Intelligent robot guidance is one of the most potential deep learning + IoT applications in the manufacturing industry. However, low costs, efficient computing, and extremely high localization accuracy are mandatory requirements for vision robot guidance, particularly in operational factories. Therefore, in this work, a low-cost edge computing-based IoT system is developed based on an innovative fine-grained attention model (FGAM). FGAM integrates a deep-learning-based attention model to detect the region of interest (ROI) and an optimized conventional computer vision model to perform fine-grained localization concentrating on the ROI. Trained with only 100 images collected from real production line, the proposed FGAM has shown superior performance over multiple benchmark models when validated using operational data. Eventually, the FGAM-based edge computing system has been deployed on a welding robot in a real-world factory for mass production. After the assembly of about 6000 products, the deployed system has achieved averaged overall process and transmission time down to 200 ms and overall localization accuracy up to 99.998%. Yinghao Chu, Daquan Feng, Zuozhu Liu, Lei Zhang 0035, Zizhou Zhao, Zhenzhong Wang, Zhiyong Feng 0001, Xiang-Gen Xia 0001 |
IEEE Internet Things J. | 3 |
| 2023 | Hierarchical Self-Supervised Learning for 3D Tooth Segmentation in Intra-Oral Mesh ScansabstractAccurately delineating individual teeth and the gingiva in the three-dimension (3D) intraoral scanned (IOS) mesh data plays a pivotal role in many digital dental applications, e.g., orthodontics. Recent research shows that deep learning based methods can achieve promising results for 3D tooth segmentation, however, most of them rely on high-quality labeled dataset which is usually of small scales as annotating IOS meshes requires intensive human efforts. In this paper, we propose a novel self-supervised learning framework, named STSNet, to boost the performance of 3D tooth segmentation leveraging on large-scale unlabeled IOS data. The framework follows two-stage training, i.e., pre-training and fine-tuning. In pre-training, three hierarchical-level, i.e., point-level, region-level, cross-level, contrastive losses are proposed for unsupervised representation learning on a set of predefined matched points from different augmented views. The pretrained segmentation backbone is further fine-tuned in a supervised manner with a small number of labeled IOS meshes. With the same amount of annotated samples, our method can achieve an mIoU of 89.88%, significantly outperforming the supervised counterparts. The performance gain becomes more remarkable when only a small amount of labeled samples are available. Furthermore, STSNet can achieve better performance with only 40% of the annotated samples as compared to the fully supervised baselines. To the best of our knowledge, we present the first attempt of unsupervised pre-training for 3D tooth segmentation, demonstrating its strong potential in reducing human efforts for annotation and verification. Zuozhu Liu, Xiaoxuan He, Hualiang Wang, Huimin Xiong, Yan Zhang 0004, Gaoang Wang, Jin Hao, Yang Feng 0011, Fudong Zhu, Haoji Hu |
IEEE Trans. Medical Imaging | 1 |
| 2022 | Renovate Yourself: Calibrating Feature Representation of Misclassified Pixels for Semantic SegmentationabstractExisting image semantic segmentation methods favor learning consistent representations by extracting long-range contextual features with the attention, multi-scale, or graph aggregation strategies. These methods usually treat the misclassified and correctly classified pixels equally, hence misleading the optimization process and causing inconsistent intra-class pixel feature representations in the embedding space during learning. In this paper, we propose the auxiliary representation calibration head (RCH), which consists of the image decoupling, prototype clustering, error calibration modules and a metric loss function, to calibrate these error-prone feature representations for better intra-class consistency and segmentation performance. RCH could be incorporated into the hidden layers, trained together with the segmentation networks, and decoupled in the inference stage without additional parameters. Experimental results show that our method could significantly boost the performance of current segmentation methods on multiple datasets (e.g., we outperform the original HRNet and OCRNet by 1.1% and 0.9% mIoU on the Cityscapes test set). Codes are available at https://github.com/VipaiLab/RCH. Hualiang Wang, Huanpeng Chu, Siming Fu, Zuozhu Liu, Haoji Hu |
AAAI | 4 |
| 2022 | Towards Calibrated Hyper-Sphere Representation via Distribution Overlap Coefficient for Long-Tailed Learning
Hualiang Wang, Siming Fu, Xiaoxuan He, Hangxiang Fang, Zuozhu Liu, Haoji Hu |
ECCV (24) | 5 |
| 2022 | Generate, Discriminate and Contrast: A Semi-Supervised Sentence Representation Learning FrameworkabstractMost sentence embedding techniques heavily rely on expensive human-annotated sentence pairs as the supervised signals.Despite the use of large-scale unlabeled data, the performance of unsupervised methods typically lags far behind that of the supervised counterparts in most downstream tasks.In this work, we propose a semi-supervised sentence embedding framework, GenSE, that effectively leverages large-scale unlabeled data.Our method include three parts: 1) Generate: A generator/discriminator model is jointly trained to synthesize sentence pairs from open-domain unlabeled corpus; 2) Discriminate: Noisy sentence pairs are filtered out by the discriminator to acquire high-quality positive and negative sentence pairs; 3) Contrast: A prompt-based contrastive approach is presented for sentence representation learning with both annotated and synthesized data.Comprehensive experiments show that GenSE achieves an average correlation score of 85.19 on the STS datasets and consistent performance improvement on four domain adaptation tasks, significantly surpassing the state-of-the-art methods and convincingly corroborating its effectiveness and generalization ability. 1 Yiming Chen 0010, Yan Zhang 0004, Bin Wang 0040, Zuozhu Liu, Haizhou Li 0001 |
EMNLP | 4 |
| 2022 | Federated Stochastic Gradient Descent Begets Self-Induced MomentumabstractFederated learning (FL) is an emerging machine learning method that can be applied in mobile edge systems, in which a server and a host of clients collaboratively train a statistical model utilizing the data and computation resources of the clients without directly exposing their privacy-sensitive data. We show that running stochastic gradient descent (SGD) in such a setting can be viewed as adding a momentum-like term to the global aggregation process. Based on this finding, we further analyze the convergence rate of a federated learning system by accounting for the effects of parameter staleness and communication resources. These results advance the understanding of the Federated SGD algorithm, and also forges a link between staleness analysis and federated computing systems, which can be useful for systems designers. Howard H. Yang, Zuozhu Liu, Yaru Fu, Tony Q. S. Quek, H. Vincent Poor |
ICASSP | 2 |
| 2022 | Energy-Efficient Intelligent Pulmonary Auscultation for Post COVID-19 Era Wearable Monitoring Enabled by Two-Stage Hybrid Neural NetworkabstractThis paper proposes an energy-efficient intelligent pulmonary auscultation system for post COVID-19 era wearable monitoring. This system consists of a tightly coupled two-stage hybrid neural network (TC-TSHNN) model and a corresponding multi-task training paradigm to improve prediction accuracy and generalization ability based on the fact that the number of COVID-19 patients is far less than that of normal people. At the first stage, two-category coarse classification is performed to identify normal and abnormal lung sounds. If the lung sound is abnormal, the second stage would be triggered to perform a four-category fine-grained classification. Besides, discrete wavelet transform is utilized for feature extraction, denoising and data reduction. In addition, advanced lightweight convolutional neural networks are used to reduce the model’s computation and improve the model’s performance. The hybrid network model can achieve 92% computation reduction and energy saving compared with a direct four-category classification when the input lung sound is normal, which is the majority of cases. Experiment results with inter-patient classification on the COVID-19 lung sound dataset from Tongji Hospital in Wuhan City and the ICBHI’17 dataset show that the proposed TC-TSHNN model can significantly reduce power consumption while maintaining competitive performance against the state-of-the-art work. Bingqiang Liu, Ziyuan Wen, Hongling Zhu, Jinsheng Lai, Jiajun Wu 0006, Heng Ping, Wenqing Liu, Guoyi Yu, Zuozhu Liu, Hesong Zeng, Chao Wang 0096 |
ISCAS | 10 |
| 2022 | Hybrid-Learning-Based Operational Visual Quality Inspection for Edge-Computing-Enabled IoT SystemabstractDeep learning-enhanced Internet of Things (IoT) plays a pivot role in advancing the transformation toward smart manufacturing, and an essential component in many smart manufacturing IoT systems is the quality inspection. However, challenges, such as expensive data labeling, innumerable types of defects, and high costs for iterative optimization, hinder the industrial applicability of previous visual surface quality inspection methods. In this article, we present an edge-computing-enabled IoT system based on an innovative hybrid learning method for visual surface quality inspection using only few labeled data and minimum iterative optimization efforts. Our hybrid learning method first employs a deep neural network to synthesize global representations of real-world industrial images, which are subsequently analyzed via an unsupervised clustering algorithm for anomaly detection. Besides, enhancement strategies, such as fine-tuning and data augmentation, are proposed to improve the robustness against the noisy data set and support low-cost inference in multiple edge devices for manufacturing operation. On a holdout data set collected from real-world factories, our method achieves classification accuracies between 90% and 98%, outperforming the benchmark method by 7%–12%. Moreover, this hybrid learning method demonstrates the effectiveness in detecting new types of surface defects and achieves test recalls between 86% and 97%, outperforming the benchmark method by 11%–34%. Yinghao Chu, Daquan Feng, Zuozhu Liu, Zizhou Zhao, Zhenzhong Wang, Xiang-Gen Xia 0001, Tony Q. S. Quek |
IEEE Internet Things J. | 3 |
| 2021 | Bootstrapped Unsupervised Sentence Representation LearningabstractYan Zhang, Ruidan He, Zuozhu Liu, Lidong Bing, Haizhou Li. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yan Zhang 0004, Ruidan He, Zuozhu Liu, Lidong Bing, Haizhou Li 0001 |
ACL/IJCNLP (1) | 3 |
| 2021 | Track without Appearance: Learn Box and Tracklet Embedding with Local and Global Motion Patterns for Vehicle TrackingabstractVehicle tracking is an essential task in the multi-object tracking (MOT) field. A distinct characteristic in vehicle tracking is that the trajectories of vehicles are fairly smooth in both the world coordinate and the image coordinate. Hence, models that capture motion consistencies are of high necessity. However, tracking with the standalone motion-based trackers is quite challenging because targets could get lost easily due to limited information, detection error and occlusion. Leveraging appearance information to assist object re-identification could resolve this challenge to some extent. However, doing so requires extra computation while appearance information is sensitive to occlusion as well. In this paper, we try to explore the significance of motion patterns for vehicle tracking without appearance information. We propose a novel approach that tackles the association issue for long-term tracking with the exclusive fully-exploited motion information. We address the tracklet embedding issue with the proposed reconstruct-to-embed strategy based on deep graph convolutional neural networks (GCN). Comprehensive experiments on the KITTI-car tracking dataset and UA-Detrac dataset show that the proposed method, though without appearance information, could achieve competitive performance with the state-of-the-art (SOTA) trackers. The source code will be available at https://github.com/GaoangW/LGMTracker. Gaoang Wang, Renshu Gu, Zuozhu Liu, Weijie Hu, Mingli Song, Jenq-Neng Hwang |
ICCV | 3 |
| 2020 | Biologically Plausible Sequence Learning with Spiking Neural Networks
Zuozhu Liu, Thiparat Chotibut, Christopher Hillar, Shaowei Lin |
AAAI | 1 |
| 2020 | Lightweight, Dynamic Graph Convolutional Networks for AMR-to-Text GenerationabstractAMR-to-text generation is used to transduce Abstract Meaning Representation structures (AMR) into text.A key challenge in this task is to efficiently learn effective graph representations.Previously, Graph Convolution Networks (GCNs) were used to encode input AMRs, however, vanilla GCNs are not able to capture non-local information and additionally, they follow a local (first-order) information aggregation scheme.To account for these issues, larger and deeper GCN models are required to capture more complex interactions.In this paper, we introduce a dynamic fusion mechanism, proposing Lightweight Dynamic Graph Convolutional Networks (LDGCNs) that capture richer non-local interactions by synthesizing higher order information from the input graphs.We further develop two novel parameter saving strategies based on the group graph convolutions and weight tied convolutions to reduce memory usage and model complexity.With the help of these strategies, we are able to train a model with fewer parameters while maintaining the model capacity.Experiments demonstrate that LDGCNs outperform stateof-the-art models on two benchmark datasets for AMR-to-text generation with significantly fewer parameters. Yan Zhang 0004, Zhijiang Guo, Zhiyang Teng, Wei Lu 0011, Shay B. Cohen, Zuozhu Liu, Lidong Bing |
EMNLP (1) | 6 |
| 2020 | An Unsupervised Sentence Embedding Method by Mutual Information MaximizationabstractBERT is inefficient for sentence-pair tasks such as clustering or semantic search as it needs to evaluate combinatorially many sentence pairs which is very time-consuming.Sentence BERT (SBERT) attempted to solve this challenge by learning semantically meaningful representations of single sentences, such that similarity comparison can be easily accessed.However, SBERT is trained on corpus with high-quality labeled sentence pairs, which limits its application to tasks where labeled data is extremely scarce.In this paper, we propose a lightweight extension on top of BERT and a novel self-supervised learning objective based on mutual information maximization strategies to derive meaningful sentence embeddings in an unsupervised manner.Unlike SBERT, our method is not restricted by the availability of labeled data, such that it can be applied on different domain-specific corpus.Experimental results show that the proposed method significantly outperforms other unsupervised sentence embedding baselines on common semantic textual similarity (STS) tasks and downstream supervised tasks.It also outperforms SBERT in a setting where in-domain labeled data is not available, and achieves performance competitive with supervised methods on various tasks. Yan Zhang 0004, Ruidan He, Zuozhu Liu, Kwan Hui Lim 0001, Lidong Bing |
EMNLP (1) | 3 |
| 2020 | Scheduling Policies for Federated Learning in Wireless NetworksabstractMotivated by the increasing computational capacity of wireless user equipments (UEs), e.g., smart phones, tablets, or vehicles, as well as the increasing concerns about sharing private data, a new machine learning model has emerged, namely federated learning (FL), that allows a decoupling of data acquisition and computation at the central unit. Unlike centralized learning taking place in a data center, FL usually operates in a wireless edge network where the communication medium is resource-constrained and unreliable. Due to limited bandwidth, only a portion of UEs can be scheduled for updates at each iteration. Due to the shared nature of the wireless medium, transmissions are subjected to interference and are not guaranteed. The performance of FL system in such a setting is not well understood. In this paper, an analytical model is developed to characterize the performance of FL in wireless networks. Particularly, tractable expressions are derived for the convergence rate of FL in a wireless setting, accounting for effects from both scheduling schemes and inter-cell interference. Using the developed analysis, the effectiveness of three different scheduling policies, i.e., random scheduling (RS), round robin (RR), and proportional fair (PF), are compared in terms of FL convergence rate. It is shown that running FL with PF outperforms RS and RR if the network is operating under a high signal-to-interference-plus-noise ratio (SINR) threshold, while RR is more preferable when the SINR threshold is low. Moreover, the FL convergence rate decreases rapidly as the SINR threshold increases, thus confirming the importance of compression and quantization of the update parameters. The analysis also reveals a trade-off between the number of scheduled UEs and subchannel bandwidth under a fixed amount of available spectrum. Howard H. Yang, Zuozhu Liu, Tony Q. S. Quek, H. Vincent Poor |
IEEE Trans. Commun. | 2 |
| 2019 | Attention-based Graph Convolutional Network for Recommendation SystemabstractMatrix completion with rating data and auxiliary information for users and items is a challenging task in recommendation systems. In this paper, we propose an end-to-end architecture named Attention-based Graph Convolutional Network (AGCN) to embed both rating data and auxiliary information in a unified space, and subsequently learn low-rank dense representations via graph convolutional networks and attention layers. Compared to previous work, AGCN reduces computational complexity with Chebyshev polynomial graph filters. The introduced attention layer, which encourages weighing the neighbor information to learn more expressive structural graph representations, can improve the prediction accuracy, and lead to faster and more stable convergence. Experimental results show that our model can perform better and converge faster than current state-of-the-art methods on the real-world MovieLens and Flixster datasets. Chenyuan Feng, Zuozhu Liu, Shaowei Lin, Tony Q. S. Quek |
ICASSP | 2 |
| 2018 | Variational Probability Flow for Biologically Plausible Training of Deep Neural NetworksabstractThe quest for biologically plausible deep learning is driven, not just by the desire to explain experimentally-observed properties of biological neural networks, but also by the hope of discovering more efficient methods for training artificial networks. In this paper, we propose a new algorithm named Variational Probably Flow (VPF), an extension of minimum probability flow for training binary Deep Boltzmann Machines (DBMs). We show that weight updates in VPF are local, depending only on the states and firing rates of the adjacent neurons. Unlike contrastive divergence, there is no need for Gibbs confabulations; and unlike backpropagation, alternating feedforward and feedback phases are not required. Moreover, the learning algorithm is effective for training DBMs with intra-layer connections between the hidden nodes. Experiments with MNIST and Fashion MNIST demonstrate that VPF learns reasonable features quickly, reconstructs corrupted images more accurately, and generates samples with a high estimated log-likelihood. Lastly, we note that, interestingly, if an asymmetric version of VPF exists, the weight updates directly explain experimental results in Spike-Timing-Dependent Plasticity (STDP). Zuozhu Liu, Tony Q. S. Quek, Shaowei Lin |
AAAI | 1 |
| 2017 | Deep fusion of heterogeneous sensor dataabstractHeterogeneous sensor data fusion is a challenging field that has gathered significant interest in recent years. In this paper, we propose a neural network-based multimodal data fusion framework named deep multimodal encoder (DME). Through our new objective function, both the intra- and inter-modal correlations of multimodal sensor data can be better exploited for recovering the missing values, and the shared representation learned can be used directly for prediction tasks. In experiments with real-world sensor data, DME shows remarkable ability for missing data imputation and new modality prediction. Compared with traditional algorithms such as kNN and Sparse-PCA, DME is more expressive, robust, and scalable to large datasets. Zuozhu Liu, Wenyu Zhang 0003, Tony Q. S. Quek, Shaowei Lin |
ICASSP | 1 |
| 2015 | Concept hierarchies and human navigationabstractWe are confronted with massive amounts of information at every turn. In order to efficiently reason about knowledge and information, humans have evolved efficient strategies for organizing complex concepts in order to form connections between and recall information. This behavior can be observed and codified when people search for objects within digital information networks. Current models of search behavior exhibit unnecessary or extraneous complexity. Minimal or simple modifications to well established algorithms yield valid models of human navigation by exploring hierarchical information inherent in networks. We explore and validate a new model of how humans navigate an information networks. To that end, we present a new path finding algorithm that approximates human navigation by leveraging the categorical classification of the nodes within the network. We compare our new model, CatPath, to existing graph distance measures when possible and show that the category paths are largely correlated with traces of human navigation. Salvador Aguiñaga, Aditya Nambiar, Zuozhu Liu, Tim Weninger |
IEEE BigData | 3 |