EDBT 2026 Demo / reviewers in the wild / expert
Mengmeng Zhang 0008
dblp:47/1678-8
· DBLP profile ↗
24ranked-venue papers
7as first author
11since 2021 · last 2026
0009-0000-7792-5878ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 6 first-author · 9 since 2021Databases, data management, data science and information retrieval · 9 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HSVC-net: hierarchical spatial-visual collaborative network for color-preserving image dehazing
Mengmeng Zhang 0008, Bowen Shao, Hongyuan Jing, Hongyun Lu |
Multim. Syst. | 1 |
| 2026 | HSV-DehazeNet: hue consistency calibration and haze density supervision for image dehazing
Hongyuan Jing, Songhao Wu, Wenlu Yang, Mengfei Han, Jinjin Hu, Kehong Li, Mengmeng Zhang 0008 |
Pattern Anal. Appl. | 8 |
| 2026 | Large-Scale Logo DetectionabstractLogo detection is crucial for trademark compliance and media monitoring, enabling companies to monitor online trademark usage and evaluate brand visibility on social media and advertisements. The use of large datasets significantly improves accuracy and generalization, emphasizing the need for high-quality datasets to optimize performance and enhance reasoning abilities in visual detection models. This drove us to create Logo4500, an unparalleled dataset featuring 4,500 logo categories and over 293,000 meticulously labeled images. To ensure the dataset's quality, we meticulously designed the construction and annotation process, with detailed information provided in our paper. Compared to existing logo datasets, Logo4500 offers greater diversity and class imbalance, making it more reflective of real-world distribution. Leveraging this high-quality dataset, we introduce a benchmark called Frequency-Aware Learnable Dual Reweighting Network (FALDR-Net), which enhances the representation of ambiguous features and addresses class imbalance for large-scale logo detection. We conducted extensive experiments, evaluating various recent methods on this new dataset and several existing publicly available logo datasets, demonstrating its effectiveness. Additionally, we verified Logo4500's generalization ability in several tasks. We anticipate that Logo4500 and the benchmark will inspire further exploration in the logo-related research community, facilitating the advancement of visual foundation models. Sujuan Hou, Weiqing Min, Jianxin Zhan, Mengmeng Zhang 0008, Peng Li 0081, Shuqiang Jiang |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | FSDG-Net: A frequency-spatial parallel network with density guidance for image dehazing
Wenlu Yang, Hongyuan Jing, Jinjin Hu, Mengfei Han, Mengmeng Zhang 0008 |
Signal Process. Image Commun. | 6 |
| 2026 | LGVF: A Low-Bitrate Generative Video Compression Framework With Spatio-Temporal Diffusion ModelabstractLow-bitrate video compression remains a fundamental challenge in video coding. Recent advances in generative technologies show great potential for addressing this problem, yet existing generative models often suffer from uncontrollability caused by error accumulation and unpredictable semantic mutation. To tackle these issues, we propose a novel low-bitrate generative video compression framework based on a spatio-temporal diffusion model. A quality comparator is designed to select the optimal feature during generation, which effectively suppresses error accumulation. In addition, we design a Semantic Mutation Feature Extraction Network (SMFN) to accurately predict frame sequences, thereby effectively handling semantically mutated features. Extensive experiments demonstrate that our framework achieves superior performance on FVD and LPIPS metrics, significantly reducing bitrate while preserving both visual fidelity and semantic consistency. Mengmeng Zhang 0008, Hongyun Lu, Huihui Bai 0001, Hongyuan Jing, Zhi Liu 0008 |
IEEE Signal Process. Lett. | 1 |
| 2026 | CondFoodGen: A Conditional Two-Stream Network for Controllable Food Image GenerationabstractFood image generation is an important research direction in food computing, aiming to produce highly realistic images that accurately capture the visual characteristics of various dishes while adhering to specified input conditions. Existing methods that rely solely on textual descriptions struggle to handle the large intra-class variability of food, often resulting in limited diversity and accuracy. Although some approaches incorporate additional conditions, they generally lack optimizations for food-specific challenges, leading to inconsistencies in texture, shape, and color fidelity. To address these limitations, we propose CondFoodGen, a diffusion-based two-stream network for controllable food image generation. The architecture consists of a control stream and a generation stream, where the control stream provides conditional guidance to regulate the generation process. To optimize bidirectional interactions between the two streams, we introduce the Bidirectional Adaptive Gating (BAG) mechanism, which not only guides synthesis but also adaptively refines control representations through feedback from the generation stream. In addition, we propose the Wavelet-Guided Hierarchical Attention (WGHA) module, which combines wavelet-based multi-frequency analysis with hierarchical attention to enhance fine-grained texture fidelity and structural realism. A progressive multi-stage training strategy further stabilizes optimization and enables seamless integration of conditional guidance with bidirectional interaction. Extensive experiments on three food image datasets demonstrate that CondFoodGen consistently generates high-quality and diverse images. Compared with the best existing food image generation methods, our approach achieves an average improvement of about 11.0% across three evaluation metrics and compared to the leading conditional generation approaches, the average improvement reaches 16.2%. The source code, trained models, and supplementary materials are publicly available at https://github.com/housujuan123/CondFoodGen. Mengyao Zhao, Hao Xiong 0001, Weiqing Min, Sujuan Hou, Mengmeng Zhang 0008, Shuqiang Jiang |
IEEE Trans. Image Process. | 5 |
| 2025 | Dual-Branch Feature Modeling and Multi-Directional Motion Perception for Video CompressionabstractEnd-to-end deep video compression in the feature space has become a key research direction, where effective feature modeling is essential. However, existing methods adopt a single-branch architecture, making it difficult to effectively perform feature modeling on complex video frames. To address this issue, we propose a Context and Fine-grained Fusion Module (CFFM), which adopts a dual-branch architecture for feature modeling. Each branch consists of two Fine-grained Spatial Aggregation Modules and two Dilated Contextual Residual Fusion Modules, focusing on local fine-grained details and global contextual information, respectively. By effectively capturing these complementary features, CFFM significantly enhances the overall feature modeling capability. Additionally, we introduce a Multi-Directional Motion Perception Module (MDMP), which integrates multi-directional and standard convolutions to dynamically expand the receptive field and adaptively model complex motion patterns, improving motion offset estimation. The entire framework is jointly optimized in an end-to-end manner. Experimental results show that our method outperforms traditional codecs and recent deep video compression approaches, achieving a 41.67% bitrate reduction compared to x265 with medium preset under PSNR, and obtaining a 3.54% gain over VVC under the MS-SSIM metric. Zhi Liu 0008, Yangbing Wang, Hongyuan Jing, Mengmeng Zhang 0008 |
MMAsia | 5 |
| 2025 | Ultra-Low Bitrate Multimodal Generative Face Video Coding Framework
Zhi Liu 0008, Hongyun Lu, Huihui Bai 0001, Hongyuan Jing, Mengmeng Zhang 0008 |
PCS | 6 |
| 2025 | GL-MambaNet: Mamba-based global and local feature fusion for image dehazing
Hongyuan Jing, Mengmeng Zhang 0008, Qiyu Rong |
Multim. Syst. | 3 |
| 2025 | An Efficient Dehazing Method Using Pixel Unshuffle and Color Correction
Hongyuan Jing, Kaiyan Wang, Aidong Chen, Mengmeng Zhang 0008 |
Signal Process. Image Commun. | 6 |
| 2025 | Enhancing Light Field Salient Object Detection With Variance-Maximized Key Focal Slice SelectionabstractLight field saliency object detection (LF SOD) methods have made significant progress recently. Most of them explore abundant multi-modal information from the all-focus image and the focal stacks at all focal planes to enrich scene details and depth perception. However, in light-field images, the spatial and depth information varies slightly across different slices, raising redundancy within focal stacks. Besides, the noise can appear repeatedly in multiple images of the focal stacks, which brings interference. To address these issues, in this work, we propose VMKNet, an effective approach that leverages innovative variance-maximized key slice selection and interacts with the all-focus image, to improve LF SOD. Specifically, we measure consistency differences between the all-focus image and each focal slice in the salient region as saliency scores. Then, we randomly assemble sets of them, where each score corresponds to a certain slice. The one exhibiting the highest variance is singled out to determine key focal slices as they reveal the diversity of salient objects. Then, the bidirectional guidance module (BGM) is presented to learn attentive features of all-focus and selected key slices in a mutual guidance manner, thus producing enhanced and holistic features. With hierarchical BGMs, our model can progressively aggregate common salient semantics and meaningful contextual details, generating more discriminative representations. Moreover, we introduce the edge enhancement module in conjunction with BGM to improve the sharpness of saliency maps. Extensive experiments on common light field datasets demonstrate that our method, termed VMKNet, outperforms recent state-of-the-art LF, RGB-D, and RGB methods. Our code is available athttps://github.com/Han-jiaxin/VMKNet. Jiaxin Han, Feng Li 0037, Mengmeng Zhang 0008, Huihui Bai 0001, Jimin Xiao, Yao Zhao 0001 |
IEEE Trans. Multim. | 4 |
| 2016 | Just Noticeable Difference Based Fast Coding Unit Partition in 3D-HEVC Intra CodingabstractSummary form only given. This paper mainly studies currently developing 3D video coding based on HEVC. HEVC-based 3D video coding mainly focuses on 3DTV and auto-stereoscopic video compression system. A variety of new encoding tools, such as inter-view motion prediction and depth modeling modes, have been added in 3D-HEVC. Although 3D-HEVC provides greater bit rate saving, it also brings the enormous encoding complexity increase. The coding time is increased correspondingly. It is necessary to reduce the encoding time. In this paper, a fast CU-sized partition algorithm is proposed for 3D-HEVC intra coding. The key point of this algorithm is to find the relationship between the texture characteristic and the sub-partition in each CU. It needs to determine whether the LCU can be subdivided to smaller CU according to the relationship. In order to reduce the redundancy of the human eye, just noticeable difference (JND) is a high efficiency model in the base of psychology and physiology. Instead of the time-consuming rate distortion optimization for coding mode decision, the variance of JND in each CU can be exploited to partition the coding unit according to human visual system characteristics. In other words, the larger blocks with higher JND variance will be subdivided to smaller blocks with lower JND variance. Consequently, the rules of CU preliminary partition are decided as follows: (a) For a 64×64 CU, if the variance of JND is larger than 0.25, the CU will be sub-divided into four 32×32 sub-blocks. (b) For a 32×32 CU, if the variance of JND is larger than 0.15, the CU will be sub-divided into four 16×16 sub-blocks. (c) For a 16×16 CU, if the variance of JND is larger than 0.10, the CU will be sub-divided into four 8×8 sub-blocks. The proposed algorithm is implemented based on HTM-13.1 reference software. The experiment condition is set up as "All Intra-Main" (AI-Main) configuration [1]. The quantization parameter (QP) values of texture are set to 25, 30, 35 and 40, respectively and the corresponding QPs of depth can be set to 34,39,42,45. The experimental results show that the fast intra mode decision algorithm provides over 29.25% encoding time saving on average with comparable rate distortion performance. Hai Ren, Huihui Bai 0001, Chunyu Lin, Mengmeng Zhang 0008, Yao Zhao 0001 |
DCC | 4 |
| 2015 | Intra-/inter-View Correlation Based Multiple Description Coding for Multiview TransmissionabstractWith the development of 3D video technology, many studies have paid attention to compression efficiency and rate distortion performance. When 3D videos are transmitted over error-prone channels, they may suffer significant quality degradation. In this paper, we combine multiview video coding (MVC) with multiple description coding (MDC) for robust transmission. The proposed scheme can give full consideration of both intra-view and inter-view correlation for better estimation. Furthermore, an adaptive mode decision is designed to generate a label as redundant information. The experiments show that the redundant information occupies just a few bits while the PSNR values of the reconstructed videos demonstrate a significant improvement. Jiansheng Guo, Huihui Bai 0001, Chunyu Lin, Mengmeng Zhang 0008, Yao Zhao 0001 |
DCC | 4 |
| 2015 | Texture Characteristics Based Fast Coding Unit Partition in HEVC Intra CodingabstractHigh efficiency video coding (HEVC) is an emerging video compression standard, developed by the Joint Collaborative Team on Video Coding (JCT-VC). The aim of HEVC standardization effort is to save about 50% bit rate for equal perceptual video quality relative to H.264/AVC. Although HEVC provides greater bit rate saving, it also brings the enormous encoding complexity increase. In this paper, we propose a fast intra CU decision algorithm based on the texture characteristics of video. Furthermore, we also consider the coding bits of each CU as auxiliary information to refine the partition results. Experimental results show that the fast intra mode decision algorithm provides over 33% complexity reduction in terms of encoding time with negligible quality loss, compared with the original HEVC test model version HM-12.0+RExt-4.0rc2. Huihui Bai 0001, Chunyu Lin, Mengmeng Zhang 0008, Yao Zhao 0001 |
DCC | 4 |
| 2014 | Two-Stage Multiview Image Compression Using Interview SIFT MatchingabstractIn this paper, a novel scheme of two-stage multiview image compression is proposed to create two-level reconstructed quality. Differently from the conventional multiview image compression algorithms, SIFT (Scale-Invariant Feature Transform) features matching from interview images are exploited to remove the correlations between multiple views. In the first stage coding, SIFT and RANSAC (RANdom SAmple Consensus) algorithms are combined to calculate the correlation matrix of interview, which then can be developed to obtain the coarse reconstruction of the current view. In the second stage coding, the reconstructed quality can be improved further by using the residual information. The experimental results have shown that at higher compression ratio, the proposed scheme can obtain better rate-distortion performance than intra coding in MVC (Multiview Video Coding). Furthermore, with the change of the compression ratio, the proposed scheme can achieve more stable reconstructed quality. Huihui Bai 0001, Mengmeng Zhang 0008, Meiqin Liu 0002, Anhong Wang, Yao Zhao 0001 |
DCC | 2 |
| 2014 | SNR Scalable Extension for 3D-HEVCabstractIn this paper we present a SNR scalable extension design on three-dimensional video compression using High Efficiency Video Coding (3D-HEVC). A multi-loop decoder solution is integrated into the proposed scalable coding serves as the whole framework for the SNR scalable 3D-HEVC. To effectively improve the coding performance, an inter-layer texture prediction is extended into the proposed scalable scenario for texture views and depth maps. To further reduce complexity and bitrate, a novel inter-layer distortion less prediction method is added because of the smooth texture characteristic in depth maps. Mengmeng Zhang 0008, Hongyun Lu, Huihui Bai 0001 |
DCC | 1 |
| 2014 | Fast Intra Prediction Based BCIM for Depth-Map in 3D-HEVCabstractThis paper presents a novel compression algorithm to replace Depth Modeling Mode for coding the depth-map. The result demonstrates that the execution time is reduced on an average 54.7% while the BD-rate of virtual views increase only 1.47%. Mengmeng Zhang 0008, Shenghui Qiu, Huihui Bai 0001 |
DCC | 1 |
| 2014 | Fast intra partition algorithm for HEVC screen content codingabstractSince the publication of the High Efficiency Video Coding standard as the newest video coding standard, several extensions have been made. Among these, the use of the screen content coding in many fields is one of the important extensions. In terms of coding tree unit (CTU) partitioning, rate distortion optimization is still used in screen content coding. The complexity of the process has resulted in problems in relation to real-time application. Thus, this paper proposes a fast-deciding CTU partition mode algorithm based on entropy and coding bits. Experimental results show that the proposed algorithm can save 32% of encoding time on average compared with the default algorithm in HM-12.1+RExt-5.1 with only 0.8% bit rate increment in coding performance. Mengmeng Zhang 0008, Yuhui Guo, Huihui Bai 0001 |
VCIP | 1 |
| 2014 | Multiple description video coding using correlation optimized temporal sampling
Huihui Bai 0001, Mengmeng Zhang 0008, Anhong Wang, Yao Zhao 0001 |
Sci. China Inf. Sci. | 2 |
| 2014 | Multiple Description Video Coding Based on Human Visual System CharacteristicsabstractIn this paper, a novel multiple description video coding scheme is proposed based on the characteristics of the human visual system (HVS). Due to the underlying spatial-temporal masking properties, human eyes cannot sense any changes below the just noticeable difference (JND) threshold. Therefore, at an encoder, only the visual information that cannot be predicted well within the JND tolerance needs to be encoded as redundant information, which leads to more effective redundancy allocation according to the HVS characteristics. Compared with the relevant existing schemes, the experimental results exhibit better performance of the proposed scheme at same bit rates, in terms of perceptual evaluation and subjective viewing. Huihui Bai 0001, Weisi Lin, Mengmeng Zhang 0008, Anhong Wang, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2013 | A fast depth-map wedgelet partitioning scheme for intra prediction in 3D video codingabstractBy employing 35 intra prediction modes, HEVC standard performs better to remove spatial redundancy between the current block and its neighbors. Although 3D video coding has adopted HEVC intra prediction and Depth Modeling Modes (DMM) technology to improve the performance, the Explicit Wedgelet Partition Mode within DMM brings unaffordable complexity to compress depth map. Based on the way of obtaining the picture texture from the mode with Sum of Absolute Transform Difference in rough mode decision, we propose a fast scheme to determine Wedgelet Partition in intra prediction, which leads to significant computational saving with marginal BD-rate increase after decoder-side view synthesis. Mengmeng Zhang 0008, Jizheng Xu, Huihui Bai 0001 |
ISCAS | 1 |
| 2012 | Multiple Description Video Coding Using Macro Block Level Correlation of Inter-/Intra-DescriptionsabstractMultiple description coding (MDC) is a promising technology for robust transmission over error-prone channels, which has attracted a lot research interests. The basic idea of MDC is to how to utilize redundant information of the descriptions for robust transmission. In view of practical applications, many MDC approaches have been proposed compatible with a certain standard codec, especially H.264/AVC. In this paper, we attempt to develop a novel MD video codec with generalized compatibility, which aims to the effective redundancy allocation from inter-/intra-descriptions. In [1], the redundancy allocation may be not enough effective due to frame level. As a result, in this paper, the redundant information will be taken into account at MB level. Huihui Bai 0001, Mengmeng Zhang 0008, Meiqin Liu 0002, Anhong Wang, Yao Zhao 0001 |
DCC | 2 |
| 2012 | Temporal Sampling Based Multiple Description Video Coding for Scenes SwitchingabstractDue to network congestion and delay sensibility, it is always a great challenge for video transmission over lossy network. Multiple description coding (MDC) is an attractive approach to solve this problem. It can efficiently combat packet loss without any retransmission thus satisfying the demand of real time services and relieving the network congestion. In view of perfect compatibility with the standard source and channel codec, temporal sampling based MDC has become a better choice for practical applications. However, for the frames switching from one scene to another temporal correlation may be destroyed by sampled in temporal domain, which may result in the false estimation when the related frames are lost at the side decoder. To address this problem, in this paper an improved MD coding based on temporal sampling is proposed to make sure the decoder can work correctly when scenes changing. Mengmeng Zhang 0008, Huihui Bai 0001 |
DCC | 1 |
| 2012 | Stereo video coding using distributed compressive sensing with joint dictionaryabstractFor many practical applications, stereo-paired video is an important special case of multiview video coding (MVC). This paper presents a novel framework of stereo video coding based on distributed compressive sensing, which also can be easily extended to MVC. According to distributed video coding (DVC) at the encoder the video sequences from each view can be compressed independently without any communications between the cameras while at the decoder the inter-view correlation can be exploited for quality enhancement. Furthermore, due to compressive sensing (CS) principles, low complexity at the encoder side can result in low power consumption in the cameras, which may be promising in wireless camera sensor network. Here, the joint dictionary is applied in compressive sensing, which can make good use of inter-view and temporal correlation for better reconstruction quality of convex optimization. The experimental results validate the effectiveness of the proposed scheme with better performance than other compared schemes. Huihui Bai 0001, Mengmeng Zhang 0008, Anhong Wang, Yao Zhao 0001 |
ICIP | 2 |