EDBT 2026 Demo / reviewers in the wild / expert
Zhenwen Liang
dblp:226/6083
· DBLP profile ↗
26ranked-venue papers
11as first author
22since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 10 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EconProver: Towards More Economical Test-Time Scaling for Automated Theorem ProvingabstractMukai Li, Linfeng Song, Zhenwen Liang, Jiahao Xu, Shansan Gong, Qi Liu, Haitao Mi, Dong Yu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Mukai Li, Linfeng Song, Zhenwen Liang, Shansan Gong, Qi Liu 0049, Haitao Mi, Dong Yu 0001 |
ACL (1) | 3 |
| 2026 | Your Reasoning Model is Secretly a Reward Model - Optimization-Free Verification from ExperienceabstractZhenwen Liang, Ruosen Li, Yujun Zhou, Linfeng Song, Dian Yu, Xinya Du, Haitao Mi, Dong Yu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhenwen Liang, Ruosen Li, Yujun Zhou 0002, Linfeng Song, Dian Yu 0001, Xinya Du, Haitao Mi, Dong Yu 0001 |
ACL (1) | 1 |
| 2026 | A Survey of Multimodal Mathematical Reasoning: From Perception, Alignment to ReasoningabstractTianyu Yang, Sihong Wu, Yilun Zhao, Zhenwen Liang, Lisen Dai, Chen Zhao, Minhao Cheng, Arman Cohan, Xiangliang Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Sihong Wu, Yilun Zhao 0001, Zhenwen Liang, Lisen Dai, Chen Zhao 0013, Minhao Cheng, Arman Cohan, Xiangliang Zhang 0001 |
ACL (1) | 4 |
| 2025 | MedAI-SciTS: Enhancing Interdisciplinary Collaboration between AI Researchers and Medical Experts
Chen Cao 0005, Zoe Xiao Fang, Zhenwen Liang, Lena Mamykina, Laura Sbaffi, Xuhai Xu |
CHI | 4 |
| 2025 | MMVU: Measuring Expert-Level Multi-Discipline Video UnderstandingabstractWe introduce $\color{Blue}{\text{MMVU}}$, a comprehensive expert-level, multi-discipline benchmark for evaluating foundation models in video understanding. $\color{Blue}{\text{MMVU}}$ includes 3,000 expert-annotated questions spanning 27 subjects across four core disciplines: Science, Healthcare, Humanities & Social Sciences, and Engineering. Compared to prior benchmarks, $\color{Blue}{\text{MMVU}}$ features three key advancements. First, it challenges models to apply domain-specific knowledge and perform expert-level reasoning to analyze specialized-domain videos, moving beyond the basic visual perception typically assessed in current video benchmarks. Second, each example is annotated by human experts from scratch. We implement strict data quality controls to ensure the high quality of the dataset. Finally, each example is enriched with expert-annotated reasoning rationals and relevant domain knowledge, facilitating in-depth analysis. We conduct an extensive evaluation of 36 frontier multimodal foundation models on $\color{Blue}{\text{MMVU}}$. The latest System-2-capable models, o1 and Gemini 2.0 Flash Thinking, achieve the highest performance among the tested models. However, they still fall short of matching human expertise. Through in-depth error analyses and case studies, we offer actionable insights for future advancements in expert-level, knowledge-intensive video understanding for specialized domains. Yilun Zhao 0001, Haowei Zhang 0002, Lujing Xie, Tongyan Hu, Guo Gan, Yitao Long, Weiyuan Chen, Chuhan Li, Chengye Wang, Ziyao Shangguan, Zhenwen Liang, Yixin Liu 0003, Chen Zhao 0013, Arman Cohan |
CVPR | 13 |
| 2025 | Learning Molecular Representation in a CellabstractPredicting drug efficacy and safety in vivo requires information on biological responses (e.g., cell morphology and gene expression) to small molecule perturbations. However, current molecular representation learning methods do not provide a comprehensive view of cell states under these perturbations and struggle to remove noise, hindering model generalization. We introduce the Information Alignment (InfoAlign) approach to learn molecular representations through the information bottleneck method in cells. We integrate molecules and cellular response data as nodes into a context graph, connecting them with weighted edges based on chemical, biological, and computational criteria. For each molecule in a training batch, InfoAlign optimizes the encoder's latent representation with a minimality objective to discard redundant structural information. A sufficiency objective decodes the representation to align with different feature spaces from the molecule's neighborhood in the context graph. We demonstrate that the proposed sufficiency objective for alignment is tighter than existing encoder-based contrastive methods. Empirically, we validate representations from InfoAlign in two downstream applications: molecular property prediction against up to 27 baseline methods across four datasets, plus zero-shot molecule-morphology matching. The code and model are available at https://github.com/liugangcode/InfoAlign. Gang Liu 0025, Srijit Seal, John Arevalo, Zhenwen Liang, Anne E. Carpenter, Meng Jiang 0001, Shantanu Singh |
ICLR | 4 |
| 2025 | MPS-Prover: Advancing Stepwise Theorem Proving by Multi-Perspective Search and Data CurationabstractAutomated Theorem Proving (ATP) in formal languages remains a formidable challenge in AI, demanding rigorous logical deduction and navigating vast search spaces. While large language models (LLMs) have shown promising performance, existing stepwise provers often suffer from biased search guidance, leading to inefficiencies and suboptimal proof strategies. This paper introduces the Multi-Perspective Search Prover (MPS-Prover), a novel stepwise ATP system designed to overcome these limitations. MPS-Prover incorporates two key innovations: a highly effective post-training data curation strategy that prunes approximately 40\% of redundant training data without sacrificing performance, and a multi-perspective tree search mechanism. This search integrates a learned critic model with strategically designed heuristic rules to diversify tactic selection, prevent getting trapped in unproductive states, and enhance search robustness. Extensive evaluations demonstrate that MPS-Prover achieves state-of-the-art performance on multiple challenging benchmarks, including miniF2F and ProofNet, outperforming prior 7B parameter models. Furthermore, our analyses reveal that MPS-Prover generates significantly shorter and more diverse proofs compared to existing stepwise and whole-proof methods, highlighting its efficiency and efficacy. Our work advances the capabilities of LLM-based formal reasoning and offers a robust framework and a comprehensive analysis for developing more powerful theorem provers. Zhenwen Liang, Linfeng Song, Tao Yang 0033, Haitao Mi, Dong Yu 0001 |
NeurIPS | 1 |
| 2024 | Chain-of-Layer: Iteratively Prompting Large Language Models for Taxonomy Induction from Limited ExamplesabstractAutomatic taxonomy induction is crucial for web search, recommendation systems, and question answering. Manual curation of taxonomies is expensive in terms of human effort, making automatic taxonomy construction highly desirable. In this work, we introduce Chain-of-Layer which is an in-context learning framework designed to induct taxonomies from a given set of entities. Chain-of-Layer breaks down the task into selecting relevant candidate entities in each layer and gradually building the taxonomy from top to bottom. To minimize errors, we introduce the Ensemble-based Ranking Filter to reduce the hallucinated content generated at each iteration. Through extensive experiments, we demonstrate that Chain-of-Layer achieves state-of-the-art performance on four real-world benchmarks. Source code available at: https://github.com/qingkaizeng/chain-of-layer. Qingkai Zeng 0001, Yuyang Bai, Zhaoxuan Tan, Shangbin Feng, Zhenwen Liang, Zhihan Zhang 0001, Meng Jiang 0001 |
CIKM | 5 |
| 2024 | MinT: Boosting Generalization in Mathematical Reasoning via Multi-view Fine-tuningabstractReasoning in mathematical domains remains a significant challenge for relatively small language models (LMs). Many current methods focus on specializing LMs in mathematical reasoning and rely heavily on distilling knowledge from powerful yet inefficient large LMs (LLMs). In this work, we explore a new direction that avoids over-reliance on LLM teachers, introducing a multi-view fine-tuning method that efficiently exploits existing mathematical problem datasets with diverse annotation styles. Our approach uniquely considers the various annotation formats as different “views” that may help each other and leverage them in training the model. By postpending distinct instructions to input questions, models can learn to generate solutions in diverse formats in a flexible manner. Experimental results show that our strategy enables relatively small LMs to outperform prior approaches that heavily rely on knowledge distillation, as well as carefully established baselines. Additionally, the proposed method grants the models promising generalization ability across various views and datasets, and the capability to learn from inaccurate or incomplete noisy data. We hope our multi-view training paradigm could inspire future studies in other machine reasoning domains. Zhenwen Liang, Dian Yu 0001, Xiaoman Pan, Wenlin Yao, Xiangliang Zhang 0001, Dong Yu 0001 |
LREC/COLING | 1 |
| 2024 | Defending Jailbreak Prompts via In-Context Adversarial GameabstractLarge Language Models (LLMs) demonstrate remarkable capabilities across diverse applications.However, concerns regarding their security, particularly the vulnerability to jailbreak attacks, persist.Drawing inspiration from adversarial training in deep learning and LLM agent learning processes, we introduce the In-Context Adversarial Game (ICAG) for defending against jailbreaks without the need for fine-tuning.ICAG leverages agent learning to conduct an adversarial game, aiming to dynamically extend knowledge to defend against jailbreaks.Unlike traditional methods that rely on static datasets, ICAG employs an iterative process to enhance both the defense and attack agents.This continuous improvement process strengthens defenses against newly generated jailbreak prompts.Our empirical studies affirm ICAG's efficacy, where LLMs safeguarded by ICAG exhibit significantly reduced jailbreak success rates across various attack scenarios.Moreover, ICAG demonstrates remarkable transferability to other LLMs, indicating its potential as a versatile defense mechanism.The code is available at https://github.com/YujunZhou/ In-Context-Adversarial-Game.28 1 82 8 8 28 3 31 9 2 1 92 82 3 31 2 98 !" "1 3"1 28 1 82 8 8 28 3 31 # $ % # &81 !" ' " 2("2 1 % 28 1 82 8 8 28 3 31 01 23 56 781 92 81 2 8 6 9 2 1 92 82 3 31 2 98 !" "1 3"1 28 1 82 8 8 28 3 31 # $ % # &81 !" ' " 2("2 1 % "&&2 !" 8 28 1 8 8 28 "&&2 !" ) "1 ! 8 2 8 28 1 8 ! 8 2) 9 *+8& 91 ,8 2 *+8& 91 ) 2'2 ! 8 2 8 28% 2 8 28 3 31 28 1 82 8 2 8 28 3 31 01 23 56 781 92 81 2 8 6 "5-.2! 2 2 "/-*+8& *+8&2 .2! 22 .2! 2 2 001 *+8& 001 .2! 2 2 * 1 81 001 (a) Self Reminder 01 23 56 781 92 81 2 8 6 28 1 82 8 8 28 3 31 9 2 1 92 82 3 31 2 98 !" "1 3"1 28 1 82 8 8 28 3 31 # $ % # &81 !" ' " 2("2 1 % 28 1 82 8 8 28 3 31 01 23 56 781 92 81 2 8 6 9 2 1 92 82 3 31 2 98 !" "1 3"1 28 1 82 8 8 28 3 31 # $ % # &81 !" ' " 2("2 1 % "&&2 !" 8 28 1 8 8 28 "&&2 !" ) "1 ! 8 2 8 28 1 8 ! 8 2) 9 *+8& 91 ,8 2 *+8& 91 ) 2'2 ! 8 2 8 28% 2 8 28 3 31 28 1 82 8 2 8 28 3 31 01 23 56 781 92 81 2 8 6 "5-.2! 2 2 "/-*+8& *+8&2 .2! 22 .2! 2 2 001 *+8& 001 .2! 2 2 * 1 81 001 (b) Our proposed In-Context Adversarial Game 0 1 2 :. 1 9 2 15 9 5 5 4-2 1*"1.9 "1 7 2 1!9 9 4 0 1 2 *"1.9 "1 7 : 2 1!9 9 4 0 1 2 : 7 :.7 18 4-9 9 -9 7 :C., 49 -99 1 =2 9 7 ::.7 49 9 ; 9 45 4 9 A49 5 :.B A49 5 :C.7 * 7 :2 1 2 1 7 :.7 18 4-9 9 -9 ,!! 57 7 : Yujun Zhou 0002, Yufei Han 0001, Haomin Zhuang, Kehan Guo, Zhenwen Liang, Hongyan Bao, Xiangliang Zhang 0001 |
EMNLP | 5 |
| 2024 | Learn Beyond The Answer: Training Language Models with Reflection for Mathematical ReasoningabstractSupervised fine-tuning enhances the problemsolving abilities of language models across various mathematical reasoning tasks.To maximize such benefits, existing research focuses on broadening the training set with various data augmentation techniques, which is effective for standard single-round question-answering settings.Our work introduces a novel technique aimed at cultivating a deeper understanding of the training problems at hand, enhancing performance not only in standard settings but also in more complex scenarios that require reflective thinking.Specifically, we propose reflective augmentation, a method that embeds problem reflection into each training instance.It trains the model to consider alternative perspectives and engage with abstractions and analogies, thereby fostering a thorough comprehension through reflective reasoning.Extensive experiments validate the achievement of our aim, underscoring the unique advantages of our method and its complementary nature relative to existing augmentation techniques. 1 Question Answer Question Zhihan Zhang 0001, Tao Ge 0001, Zhenwen Liang, Wenhao Yu 0002, Dian Yu 0001, Mengzhao Jia, Dong Yu 0001, Meng Jiang 0001 |
EMNLP | 3 |
| 2024 | Can LLMs Solve Molecule Puzzles? A Multimodal Benchmark for Molecular Structure ElucidationabstractLarge Language Models (LLMs) have shown significant problem-solving capabilities across predictive and generative tasks in chemistry. However, their proficiency in multi-step chemical reasoning remains underexplored. We introduce a new challenge: molecular structure elucidation, which involves deducing a molecule’s structure from various types of spectral data. Solving such a molecular puzzle, akin to solving crossword puzzles, poses reasoning challenges that require integrating clues from diverse sources and engaging in iterative hypothesis testing. To address this challenging problem with LLMs, we present \textbf{MolPuzzle}, a benchmark comprising 217 instances of structure elucidation, which feature over 23,000 QA samples presented in a sequential puzzle-solving process, involving three interlinked sub-tasks: molecule understanding, spectrum interpretation, and molecule construction. Our evaluation of 12 LLMs reveals that the best-performing LLM, GPT-4o, performs significantly worse than humans, with only a small portion (1.4\%) of its answers exactly matching the ground truth. However, it performs nearly perfectly in the first subtask of molecule understanding, achieving accuracy close to 100\%. This discrepancy highlights the potential of developing advanced LLMs with improved chemical reasoning capabilities in the other two sub-tasks. Our MolPuzzle dataset and evaluation code are available at this \href{https://github.com/KehanGuo2/MolPuzzle}{link}. Kehan Guo, Bozhao Nan, Yujun Zhou 0002, Taicheng Guo, Zhichun Guo, Mihir Surve, Zhenwen Liang, Nitesh V. Chawla, Olaf Wiest, Xiangliang Zhang 0001 |
NeurIPS | 7 |
| 2023 | Generalizing Math Word Problem Solvers via Solution DiversificationabstractCurrent math word problem (MWP) solvers are usually Seq2Seq models trained by the (one-problem; one-solution) pairs, each of which is made of a problem description and a solution showing reasoning flow to get the correct answer. However, one MWP problem naturally has multiple solution equations. The training of an MWP solver with (one-problem; one-solution) pairs excludes other correct solutions, and thus limits the generalizability of the MWP solver. One feasible solution to this limitation is to augment multiple solutions to a given problem. However, it is difficult to collect diverse and accurate augment solutions through human efforts. In this paper, we design a new training framework for an MWP solver by introducing a solution buffer and a solution discriminator. The buffer includes solutions generated by an MWP solver to encourage the training data diversity. The discriminator controls the quality of buffered solutions to participate in training. Our framework is flexibly applicable to a wide setting of fully, semi-weakly and weakly supervised training for all Seq2Seq MWP solvers. We conduct extensive experiments on a benchmark dataset Math23k and a new dataset named Weak12k, and show that our framework improves the performance of various MWP solvers under different settings by generating correct and diverse solutions. Zhenwen Liang, Lei Wang 0185, Yan Wang 0060, Jie Shao 0001, Xiangliang Zhang 0001 |
AAAI | 1 |
| 2023 | Let GPT be a Math Tutor: Teaching Math Word Problem Solvers with Customized Exercise GenerationabstractIn this paper, we present a novel approach for distilling math word problem solving capabilities from large language models (LLMs) into smaller, more efficient student models. Our approach is designed to consider the student model's weaknesses and foster a tailored learning experience by generating targeted exercises aligned with educational science principles, such as knowledge tracing and personalized learning. Concretely, we let GPT-3 be a math tutor and run two steps iteratively: 1) assessing the student model's current learning status on a GPT-generated exercise book, and 2) improving the student model by training it with tailored exercise samples generated by GPT-3. Experimental results reveal that our approach outperforms LLMs (e.g., GPT-3 and PaLM) in accuracy across three distinct benchmarks while employing significantly fewer parameters. Furthermore, we provide a comprehensive analysis of the various components within our methodology to substantiate their efficacy. Zhenwen Liang, Wenhao Yu 0002, Tanmay Rajpurohit, Peter Clark, Xiangliang Zhang 0001, Ashwin Kalyan |
EMNLP | 1 |
| 2023 | UniMath: A Foundational and Multimodal Mathematical ReasonerabstractWhile significant progress has been made in natural language processing (NLP), existing methods exhibit limitations in effectively interpreting and processing diverse mathematical modalities.Therefore, we introduce UniMath, a versatile and unified system designed for multimodal mathematical reasoning tasks.Tackling complex problem-solving in arithmetic, geometry, and table-based math, UniMath utilizes a fine-tuned T5 model augmented with a variational autoencoder (VAE)-based image tokenizer.By jointly training and evaluating the model on three diverse datasets -SVAMP, GeoQA, and TableMWP, UniMath achieves state-of-the-art performance.The model's generalization ability is further demonstrated via fine-tuning on two additional datasets, MathQA and Geo-Proving.Through comprehensive evaluations, we show that joint training across diverse math tasks improves overall model performance and enhances its ability to generalize across different mathematical reasoning tasks.This pioneering approach provides a blueprint and inspires further efforts on unified mathematical reasoning with deep learning systems. Zhenwen Liang, Xiangliang Zhang 0001 |
EMNLP | 1 |
| 2023 | Don't be Blind to Questions: Question-Oriented Math Word Problem SolvingabstractZhenwen Liang, Jipeng Zhang, Xiangliang Zhang. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Zhenwen Liang, Xiangliang Zhang 0001 |
IJCNLP (1) | 1 |
| 2023 | What can Large Language Models do in chemistry? A comprehensive benchmark on eight tasksabstractLarge Language Models (LLMs) with strong abilities in natural language processing tasks have emerged and have been applied in various kinds of areas such as science, finance and software engineering. However, the capability of LLMs to advance the field of chemistry remains unclear. In this paper, rather than pursuing state-of-the-art performance, we aim to evaluate capabilities of LLMs in a wide range of tasks across the chemistry domain. We identify three key chemistry-related capabilities including understanding, reasoning and explaining to explore in LLMs and establish a benchmark containing eight chemistry tasks. Our analysis draws on widely recognized datasets facilitating a broad exploration of the capacities of LLMs within the context of practical chemistry. Five LLMs (GPT-4,GPT-3.5, Davinci-003, Llama and Galactica) are evaluated for each chemistry task in zero-shot and few-shot in-context learning settings with carefully selected demonstration examples and specially crafted prompts. Our investigation found that GPT-4 outperformed other models and LLMs exhibit different competitive levels in eight chemistry tasks. In addition to the key findings from the comprehensive benchmark analysis, our work provides insights into the limitation of current LLMs and the impact of in-context learning settings on LLMs’ performance across various chemistry tasks. The code and datasets used in this study are available at https://github.com/ChemFoundationModels/ChemLLMBench. Taicheng Guo, Kehan Guo, Bozhao Nan, Zhenwen Liang, Zhichun Guo, Nitesh V. Chawla, Olaf Wiest, Xiangliang Zhang 0001 |
NeurIPS | 4 |
| 2022 | Analogical Math Word Problems Solving with Enhanced Problem-Solution AssociationabstractMath word problem (MWP) solving is an important task in question answering which requires human-like reasoning ability.Analogical reasoning has long been used in mathematical education, as it enables students to apply common relational structures of mathematical situations to solve new problems.In this paper, we propose to build a novel MWP solver by leveraging analogical MWPs, which advance the solver's generalization ability across different kinds of MWPs.The key idea, named analogy identification, is to associate the analogical MWP pairs in a latent space, i.e., encoding an MWP close to another analogical MWP, while moving away from the non-analogical ones.Moreover, a solution discriminator is integrated into the MWP solver to enhance the association between the representations of MWPs and their true solutions.The evaluation results verify that our proposed analogical learning strategy promotes the performance of MWP-BERT on Math23k over the state-of-theart model Generate2Rank, with 5 times fewer parameters in the encoder.We also find that our model has a stronger generalization ability in solving difficult MWPs due to the analogical learning from easy MWPs. Zhenwen Liang, Xiangliang Zhang 0001 |
EMNLP | 1 |
| 2022 | ArMATH: a Dataset for Solving Arabic Math Word ProblemsabstractThis paper studies solving Arabic Math Word Problems by deep learning. A Math Word Problem (MWP) is a text description of a mathematical problem that can be solved by deriving a math equation to reach the answer. Effective models have been developed for solving MWPs in English and Chinese. However, Arabic MWPs are rarely studied. This paper contributes the first large-scale dataset for Arabic MWPs, which contains 6,000 samples of primary-school math problems, written in Modern Standard Arabic (MSA). Arabic MWP solvers are then built with deep learning models and evaluated on this dataset. In addition, a transfer learning model is built to let the high-resource Chinese MWP solver promote the performance of the low-resource Arabic MWP solver. This work is the first to use deep learning methods to solve Arabic MWP and the first to use transfer learning to solve MWP across different languages. The transfer learning enhanced solver has an accuracy of 74.15%, which is 3% higher than the solver without using transfer learning. We make the dataset and solvers available in public for encouraging more research of Arabic MWPs: https://github.com/reem-codes/ArMATH Reem Alghamdi, Zhenwen Liang, Xiangliang Zhang 0001 |
LREC | 2 |
| 2021 | Solving Math Word Problems with Teacher SupervisionabstractMath word problems (MWPs) have been recently addressed with Seq2Seq models by `translating' math problems described in natural language to a mathematical expression, following a typical encoder-decoder structure. Although effective in solving classical math problems, these models fail when a subtle variation is applied to the word expression of a math problem, and leads to a remarkably different answer. We find the failure is because MWPs with different answers but similar math formula expression are encoded closely in the latent space. We thus designed a teacher module to make the MWP encoding vector match the correct solution and disaccord from the wrong solutions, which are manipulated from the correct solution. Experimental results on two benchmark MWPs datasets verified that our proposed solution outperforms the state-of-the-art models. Zhenwen Liang, Xiangliang Zhang 0001 |
IJCAI | 1 |
| 2021 | Multi-Branch Networks for Video Super-Resolution With Dynamic Reconstruction StrategyabstractRecently, the rapid development of 2-dimensional (2D) convolutional neural networks (CNNs) has driven single image super-resolution (SISR) into a new era, owing to their powerful ability in modeling spatial relation within one single image. However, few studies focus on video super-resolution (VSR) due to the key challenge that apart from the spatial relation, the temporal dependence among consecutive low-resolution (LR) frames must be taken into consideration for better reconstruction. In this article, unlike most previous methods based on optical flow for motion compensation, 3-dimensional (3D) convolution is utilized to capture the temporal relation. Firstly, in contrast to the conventional 3D convolution which is notorious for the excessively high computational burden, we propose an efficient 3D convolutional block (E3DB) through convolution factorization principle (CFP), which significantly reduces the computing load while maximally maintaining the temporal information. Then, by taking advantage of E3DB, we propose a novel multi-resolution extraction block (MREB) which aggregates the information from multiple resolutions, leading to the stronger high-resolution representation learning and better feature extraction. Besides, based on our 3-branch architecture, instead of the simple addition or concatenation, a dynamic reconstruction strategy (DRS) is proposed to adaptively fuse the optimal information of temporal dependence from each branch. It is therefore termed as dynamic multiple branch network (DMBN). Comprehensive experiments on public benchmark datasets demonstrate the superiority of our DMBN over the current state-of-the-art methods in terms of accuracy and efficiency. Dongyang Zhang 0001, Jie Shao 0001, Zhenwen Liang, Xueliang Liu, Heng Tao Shen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Large Factor Image Super-Resolution With Cascaded Convolutional Neural NetworksabstractRecently, convolutional neural networks (CNNs) have attracted considerable attention in single image super-resolution (SISR) and have enabled great performance improvements. However, most of the existing methods super-resolve input images to the desired size with an interpolation operation during the beginning stage, which brings about heavy aliasing artifacts and high computational costs. Especially for large upsampling factors (e.g., 8×), it remains a challenge to restore high-quality results for deeply degraded images. To tackle this problem, we propose a cascaded super-resolution convolutional neural network (CSRCNN), which takes a single low-resolution (LR) image as an input and reconstructs high-resolution (HR) images in a progressive way. At each cascaded level, to help converge and improve the accuracy, a novel U-net based block with backprojection is first introduced, which exploits the mutual relation between HR and LR feature spaces. A refined block following the U-net block is also used to reconstruct the realistic texture details. In addition, we naturally utilize the strategy of curriculum learning, organizing the learning process from easy (small factors) to hard (large factors). Comprehensive experiments on benchmark datasets demonstrate that the proposed network achieves superior results compared with those of other state-of-the-art methods, particularly with the 8× upsampling factor. Dongyang Zhang 0001, Jie Shao 0001, Zhenwen Liang, Lianli Gao, Heng Tao Shen |
IEEE Trans. Multim. | 3 |
| 2020 | Fused Recurrent Network Via Channel Attention For Remote Sensing Satellite Image Super-ResolutionabstractRemote sensing satellite images often suffer from low spatial resolution. Image super-resolution plays an important role in remote sensing image processing. However, existing methods show that increasing network depth will inevitably lead to the dramatic increase of model parameters and the over-fitting problem. Besides, most methods treat different types of information (low-frequency and high-frequency) equally. Motivated by these observations, we propose a fused recurrent network via channel attention (CA-FRN) in this paper. The basic module, recursive channel attention block (RCAB), pays enough attention to the high-frequency information and diminishes the low-frequency information adaptively through channel attention. Based on RCAB, we render our model effective by retaining and fusing hierarchical local information of both low-resolution and high-resolution, and we enhance the network performance simply by increasing the number of RCABs without adding extra parameters. We evaluate the proposed model on satellite images from different datasets, and the proposed CA-FRN is superior to the state-of-the-art methods. Code is available at https://github.com/lxy0922/CAFRN. Dongyang Zhang 0001, Zhenwen Liang, Deqiang Ouyang, Jie Shao 0001 |
ICME | 3 |
| 2020 | Joint image deblurring and super-resolution with attention dual supervised network
Dongyang Zhang 0001, Zhenwen Liang, Jie Shao 0001 |
Neurocomputing | 2 |
| 2020 | Traffic sign detection and recognition based on pyramidal convolutional networks
Zhenwen Liang, Jie Shao 0001, Dongyang Zhang 0001, Lianli Gao |
Neural Comput. Appl. | 1 |
| 2019 | Jointly Solving Deblurring and Super-Resolution Problems with Dual Supervised NetworkabstractThanks to the vigorous development of deep learning techniques, the solutions of super-resolution (SR) and deblurring have made great progress in recent years. However, in real-world scenarios, the low-quality images not only suffer from the loss of spatial information (resolution), they are often degradated by complicated non-uniform motion blur as well. As a result, traditional SR and deblurring methods are not robust enough to restore those blurred and low-resolution (LR) images to sharp and high-resolution (HR) images. Therefore, we propose a novel dual supervised network (DSN) to jointly solve the SR and deblurring problems. In the beginning we use a deblurring module to make blurry input much sharper and feed the deblurred feature to a reconstruction module, which aims to produce the final HR result from the deblurred feature. Moreover, we employ dual supervised structure to fully exploit the mutual dependencies between LR and HR patches. Experiments show that our network outperforms the state-of-the-art methods. Zhenwen Liang, Dongyang Zhang 0001, Jie Shao 0001 |
ICME | 1 |