EDBT 2026 Demo / reviewers in the wild / expert
Xuan Dong 0001
dblp:86/1293-1
· DBLP profile ↗
25ranked-venue papers
15as first author
12since 2021 · last 2026
0000-0002-3121-9756ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 10 first-author · 7 since 2021Artificial intelligence and machine learning · 11 · 7 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Temporal Group Constrained Transformer With Deformable Landmark Attention for Video Dimensional Emotion RecognitionabstractVideo dimensional emotion recognition aims to map human affect into the dimensional emotion space based on visual signals. Recent works notice that it is beneficial to locate key facial regions related to human emotion perception, as well as establish long-term temporal dependencies. While preliminary attempts have been made, there still exists much space for further improvements. In this paper, to better exploit key facial regions, we propose the Temporal cue guided Deformable Landmark Spatial (TDLS) transformer which attends to key facial regions in a data-dependent manner. We also propose the temporal cue guided frame representation learning to learn the spatial representation of each frame by considering features of other frames together. To better model temporal dependencies, we propose the Multi-layer Group Constrained Temporal (MGCT) transformer to summarize features of frames to multi-layer groups, perform group-to-group communications, and let group-level features guide the frame-level emotion recognition. We also introduce cross-clip representation learning to generate consistent results across different clips and videos. Extensive experiments are conducted on two benchmark datasets and superior results are achieved by our method compared to state-of-the-art approaches. Weixin Li 0001, Xiangjing Meng, Linmei Hu, Xuan Dong 0001 |
IEEE Trans. Affect. Comput. | 4 |
| 2025 | KDA: Knowledge Distillation Adversarial Framework With Vision Foundation Models for Landslide SegmentationabstractLandslides pose severe threats to infrastructure and safety, and their segmentation in remote sensing imagery remains challenging due to irregular boundaries, scale variation, and complex terrain. Traditional lightweight models often struggle to capture rich semantic features under these conditions. To address this, we leverage vision foundation models (VFMs) as teachers and propose a knowledge distillation adversarial (KDA) framework to transfer high-capacity knowledge into compact student models. Additionally, we introduce a dynamic cross-layer fusion (DCF) decoder to enhance global–local feature interaction. The experimental results demonstrate that, compared to the previous best-performing model SegNeXt [89.92% precision and 84.78% mean intersection over union (mIoU)], our method achieves a precision of 91.93% and mIoU of 86.53%, yielding improvements of 2.01% and 1.75%, respectively. Source code is available athttps://github.com/PreWisdom/KDA Lulin Li, Xuan Dong 0001, Lei Shi 0002, Pin Tao |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2025 | Video Demoireing Using Focused-Defocused Dual-Camera SystemabstractMoire patterns, unwanted color artifacts in images and videos, arise from the interference between spatially high-frequency scene contents and the spatial discrete sampling of digital cameras. Existing demoireing methods primarily rely on single-camera image/video processing, which faces two critical challenges: 1) distinguishing moire patterns from visually similar real textures, and 2) preserving tonal consistency and temporal coherence while removing moire artifacts. To address these issues, we propose a dual-camera framework that captures synchronized videos of the same scene: one in focus (retaining high-quality textures but may exhibit moire patterns) and one defocused (with significantly reduced moire patterns but blurred textures). We use the defocused video to help distinguish moire patterns from real texture, so as to guide the demoireing of the focused video. We propose a frame-wise demoireing pipeline, which begins with an optical flow based alignment step to address any discrepancies in displacement and occlusion between the focused and defocused frames. Then, we leverage the aligned defocused frame to guide the demoireing of the focused frame using a multi-scale CNN and a multi-dimensional training loss. To maintain tonal and temporal consistency, our final step involves a joint bilateral filter to leverage the demoireing result from the CNN as the guide to filter the input focused frame to obtain the final output. Experimental results demonstrate that our proposed framework largely outperforms state-of-the-art image and video demoireing methods. Xuan Dong 0001, Xiangyuan Sun, Ya Li 0001, Weixin Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | ReWiTe: Realistic Wide-angle and Telephoto Dual Camera Fusion Dataset via Beam Splitter Camera RigabstractThe fusion of images from dual camera systems featuring a wide-angle and a telephoto camera has become a hotspot problem recently. By integrating simultaneously captured wide-angle and telephoto images from these systems, the resulting fused image achieves a wide field of view (FOV) coupled with high-definition quality. Existing approaches are mostly deep learning methods, and predominantly rely on supervised learning, where the training dataset plays a pivotal role. However, current datasets typically adopt a data synthesis approach, where the wide-angle inputs are synthesized rather than captured using real wide-angle cameras, and the ground-truth image is captured by wide-angle cameras whose quality is substantially lower than that of input telephoto images captured by telephoto cameras. To address these limitations, we introduce a novel hardware setup utilizing a beam splitter to simultaneously capture three images, i.e. input pairs and ground-truth images, from two authentic cellphones equipped with wide-angle and telephoto dual cameras. Specifically, the wide-angle and telephoto images captured by cellphone 2 serve as the input pair, while the telephoto image captured by cellphone 1, which is calibrated to match the optical path of the wide-angle image from cellphone 2, serves as the ground-truth image, maintaining quality on par with the input telephoto image. Experiments validate the efficacy of our newly introduced dataset, named ReWiTe, which can significantly enhance the performance of various existing methods for the real-world wide-angle and telephoto dual image fusion task. Chunli Peng, Xuan Dong 0001, Zhengqing Li, Weixin Li 0001 |
ACM Multimedia | 2 |
| 2023 | Human Emotion Recognition With Relational Region-Level AnalysisabstractRecognizing the emotional state of a person within the image in real-world scenarios is a key problem in affective computing and has various promising applications. Local regions in the image, including different objects in the background scene and parts within the foreground body, usually have different contributions to emotion perception of the target person. This, however, has not been well exploited in most existing methods. In this article, we propose to make relational region-level analysis to account for the different contributions of different regions to emotion recognition. For the background scene, we propose a Body-Object Attention (BOA) module to estimate the contributions of background objects to emotion recognition given the target foreground body. Within the foreground body, we propose a Body Part Attention (BPA) module to recalibrate the channel-wise body feature responses to attend on body parts that are more important. Moreover, we propose to model the emotion label dependency in real-world images, considering both the semantic meanings of these labels and their co-occurrence patterns. We evaluate the proposed method on the EMOTIC and CAER-S datasets, and experimental results show the superiority of our method compared with the state-of-the-art algorithms. Weixin Li 0001, Xuan Dong 0001, Yunhong Wang 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2023 | Dual-Lens HDR using Guided 3D Exposure CNN and Guided Denoising TransformerabstractWe study the high dynamic range (HDR) imaging problem in dual-lens systems. Existing methods usually treat the HDR imaging problem as an image fusion problem and the HDR result is estimated by fusing the aligned short exposure image and long exposure image. However, the image fusion pipeline depends highly on the image alignment, which is difficult to be perfect. We propose to transfer the dual-lens HDR imaging problem into the disentangled enhancement of exposure correction and denoising for the short exposure image, guided by the long exposure image. In the guided exposure correction module, we make use of the guidance image and 3D color transformation to propose a guided 3D exposure CNN (GEC) to get the rough HDR result from the short exposure image. Then, in the guided denoising module, we make use of the cross-attention mechanism to propose a guided denoising transformer (GDT) to directly use the long exposure image as guidance to denoise the rough HDR result in a pyramid way. And in both modules, we bypass the difficult image alignment processing. Experimental results demonstrate the superiority of our method over the state-of-the-art ones. Weixin Li 0001, Chang Liu 0071, Xue Tian, Ya Li 0001, Xiaojie Wang 0006, Xuan Dong 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2023 | CUR Transformer: A Convolutional Unbiased Regional Transformer for Image DenoisingabstractImage denoising is a fundamental problem in computer vision and multimedia computation. Non-local filters are effective for image denoising. But existing deep learning methods that use non-local computation structures are mostly designed for high-level tasks, and global self-attention is usually adopted. For the task of image denoising, they have high computational complexity and have a lot of redundant computation of uncorrelated pixels. To solve this problem and combine the marvelous advantages of non-local filter and deep learning, we propose a Convolutional Unbiased Regional (CUR) transformer. Based on the prior that, for each pixel, its similar pixels are usually spatially close, our insights are that (1) we partition the image into non-overlapped windows and perform regional self-attention to reduce the search range of each pixel, and (2) we encourage pixels across different windows to communicate with each other. Based on our insights, the CUR transformer is cascaded by a series of convolutional regional self-attention (CRSA) blocks with U-style short connections. In each CRSA block, we use convolutional layers to extract the query, key, and value features, namely Q , K , and V , of the input feature. Then, we partition the Q , K , and V features into local non-overlapped windows and perform regional self-attention within each window to obtain the output feature of this CRSA block. Among different CRSA blocks, we perform the unbiased window partition by changing the partition positions of the windows. Experimental results show that the CUR transformer outperforms the state-of-the-art methods significantly on four low-level vision tasks, including real and synthetic image denoising, JPEG compression artifact reduction, and low-light image enhancement. Weixin Li 0001, Xiaoyan Hu 0006, Xiaojie Wang 0006, Xuan Dong 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2022 | Spatially Consistent Transformer for Colorization in Monochrome-Color Dual-Lens SystemabstractWe study the colorization problem in monochrome-color dual-lens camera systems, i.e. colorizing the gray image from the monochrome camera using the color image from the color camera as reference. In related methods, cost volume based CNN methods achieve the state-of-the-art results, but they are costly in GPU memory due to building the 4D cost volume. Recently, some slice-wise cross-attention based methods are proposed for related problems. The slice-wise cross-attention has much less costs in GPU memory but directly using them for this colorization problem cannot generate competing results. We make use of the non-local computation property of cross-attention to propose a transformer based method. To overcome the limitations of straight-forward slice-wise cross-attention, we propose the spatially consistent cross-attention (SCCA) block to encourage pixels of slices across different epipolar lines in the gray image to find spatially consistent correspondence with pixels of the reference color image. And, to further reduce the memory cost while keeping the colorization accuracy, we design a pyramid processing strategy to cascade a series of SCCA blocks with smaller slice size and perform the colorization from coarse to fine. To extract more powerful image features, we use several regional self-attention (RSA) blocks with U-style connections. Experimental results show that we outperform the state-of-the-art methods largely on the synthesized datasets of Cityscapes, Sintel, and SceneFlow, and the real monochrome-color dual-lens dataset. Xuan Dong 0001, Chang Liu 0071, Xiaoyan Hu 0006, Weixin Li 0001 |
IEEE Trans. Image Process. | 1 |
| 2022 | A Colorization Framework for Monochrome-Color Dual-Lens Systems Using a Deep Convolutional NetworkabstractIn monochrome-color dual-lens systems, the monochrome camera can capture images with higher quality than the color camera. To obtain high quality color images, a better approach is to colorize the gray images from the monochrome camera with the color images from the color camera serving as a reference. In addition, the colorization may fail in some cases, which makes the estimation of the colorization quality a necessary step before outputting the colorization result. To solve these problems, we propose a deep convolutional network based framework. 1) In the colorization module, the proposed colorization CNN uses deep feature representations, attention operation, 3-D regulation and color correction to make use of colors of multiple pixels in the reference image for colorizing each pixel in the input gray image. 2) In the colorization quality estimation module, based on the symmetry property of colorization, we propose to utilize the colorization CNN again to colorize the gray map of the original reference color image using the first-time colorization result from the colorization module as reference. Then, the quality loss of the second-time colorization result can be used for estimating the colorization quality. Experimental results show that our method can largely outperform the state-of-the-art colorization methods and estimate the colorization quality accurately as well. Xuan Dong 0001, Weixin Li 0001, Xiaoyan Hu 0006, Xiaojie Wang 0006, Yunhong Wang 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2021 | MIEHDR CNN: Main Image Enhancement based Ghost-Free High Dynamic Range Imaging using Dual-Lens SystemsabstractWe study the High Dynamic Range (HDR) imaging problem using two Low Dynamic Range (LDR) images that are shot from dual-lens systems in a single shot time with different exposures. In most of the related HDR imaging methods, the problem is usually solved by Multiple Images Merging, i.e. the final HDR image is fused from pixels of all the input LDR images. However, ghost artifacts can be hardly avoided using this strategy. Instead of directly merging the multiple LDR inputs, we use an indirect way which enhances the main image, i.e. the short exposure image IS, using the long exposure image IL serving as guidance. In detail, we propose a new model, named MIEHDR CNN model, which consists of three subnets, i.e. Soft Warp CNN, 3D Guided Denoising CNN and Fusion CNN. The Soft Warp CNN aligns IL to get the aligned result ILA using the soft exposed result of IS as reference. The 3D Guided Denoising CNN denoises the soft exposed result of IS using ILA as guidance, whose result are fed into the Fusion CNN with IS to get the HDR result. The MIEHDR CNN model is implemented by MindSpore and experimental results show that we can outperform related methods largely and avoid ghost artifacts. Xuan Dong 0001, Xiaoyan Hu 0006, Weixin Li 0001, Xiaojie Wang 0006, Yunhong Wang 0001 |
AAAI | 1 |
| 2021 | Pyramid convolutional network for colorization in monochrome-color multi-lens camera system
Xuan Dong 0001, Weixin Li 0001, Xiaojie Wang 0006 |
Neurocomputing | 1 |
| 2021 | Self-Supervised Colorization Towards Monochrome-Color Camera Systems Using Cycle CNNabstractColorization in monochrome-color camera systems aims to colorize the gray image IGfrom the monochrome camera using the color image RCfrom the color camera as reference. Since monochrome cameras have better imaging quality than color cameras, the colorization can help obtain higher quality color images. Related learning based methods usually simulate the monochrome-color camera systems to generate the synthesized data for training, due to the lack of ground-truth color information of the gray image in the real data. However, the methods that are trained relying on the synthesized data may get poor results when colorizing real data, because the synthesized data may deviate from the real data. We present a self-supervised CNN model, named Cycle CNN, which can directly use the real data from monochrome-color camera systems for training. In detail, we use the Weighted Average Colorization (WAC) network to do the colorization twice. First, we colorize IGusing RCas reference to obtain the first-time colorization result IC. Second, we colorize the de-colored map of RC, i.e. RG, using the concatenated image of IGand Cb/Cr channels of the first-time colorization result IC, i.e. ICCband ICCr, as reference to obtain the second-time colorization result RC'. In this way, for the second-time colorization result RC', we use the Cb and Cr channels of the original color map RCas ground-truth and introduce the cycle consistency loss to push RC'Cb/Cr≈ RCCb/Cr. Also, for the Y channel of the first-time colorization result ICY, we propose the Global Curve Adjustment (GCA) network and the structure similarity loss to encourage the structure similarity between ICYand IG. In addition, we introduce a spatial smoothness loss within the WAC network to encourage spatial smoothness of the colorization result. Combining all these losses, we could train the Cycle CNN using the real data in the absence of the ground-truth color information of IG. Experimental results show that we can outperform related methods largely for colorizing real data. Xuan Dong 0001, Chang Liu 0071, Weixin Li 0001, Xiaoyan Hu 0006, Xiaojie Wang 0006, Yunhong Wang 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | Cycle-CNN for Colorization towards Real Monochrome-Color Camera SystemsabstractColorization in monochrome-color camera systems aims to colorize the gray image IG from the monochrome camera using the color image RC from the color camera as reference. Since monochrome cameras have better imaging quality than color cameras, the colorization can help obtain higher quality color images. Related learning based methods usually simulate the monochrome-color camera systems to generate the synthesized data for training, due to the lack of ground-truth color information of the gray image in the real data. However, the methods that are trained relying on the synthesized data may get poor results when colorizing real data, because the synthesized data may deviate from the real data. We present a new CNN model, named cycle CNN, which can directly use the real data from monochrome-color camera systems for training. In detail, we use the colorization CNN model to do the colorization twice. First, we colorize IG using RC as reference to obtain the first-time colorization result IC. Second, we colorize the de-colored map of RC, i.e. RG, using the first-time colorization result IC as reference to obtain the second-time colorization result R′C. In this way, for the second-time colorization result R′C, we use the original color map RC as ground-truth and introduce the cycle consistency loss to push R′C ≈ RC. Also, for the first-time colorization result IC, we propose a structure similarity loss to encourage the luminance maps between IG and IC to have similar structures. In addition, we introduce a spatial smoothness loss within the colorization CNN model to encourage spatial smoothness of the colorization result. Combining all these losses, we could train the colorization CNN model using the real data in the absence of the ground-truth color information of IG. Experimental results show that we can outperform related methods largely for colorizing real data. Xuan Dong 0001, Weixin Li 0001, Xiaojie Wang 0006, Yunhong Wang 0001 |
AAAI | 1 |
| 2019 | Learning a Deep Convolutional Network for Colorization in Monochrome-Color Dual-Lens SystemabstractIn the monochrome-color dual-lens system, the gray image captured by the monochrome camera has better quality than the color image from the color camera, but does not have color information. To get high-quality color images, it is desired to colorize the gray image with the color image as reference. Related works usually use hand-crafted methods to search for the best-matching pixel in the reference image for each pixel in the input gray image, and copy the color of the best-matching pixel as the result. We propose a novel deep convolution network to solve the colorization problem in an end-to-end way. Based on our observation that, for each pixel in the input image, there usually exist multiple pixels in the reference image that have the correct colors, our method performs weighted average of colors of the candidate pixels in the reference image to utilize more candidate pixels with correct colors. The weight values between pixels in the input image and the reference image are obtained by learning a weight volume using deep feature representations, where an attention operation is proposed to focus on more useful candidate pixels and a 3-D regulation is performed to learn with context information. In addition, to correct wrongly colorized pixels in occlusion regions, we propose a color residue joint learning module to correct the colorization result with the input gray image as guidance. We evaluate our method on the Scene Flow, Cityscapes, Middlebury, and Sintel datasets. Experimental results show that our method largely outperforms the state-of-the-art methods. Xuan Dong 0001, Weixin Li 0001, Xiaojie Wang 0006, Yunhong Wang 0001 |
AAAI | 1 |
| 2019 | Shoot high-quality color images using dual-lens system with monochrome and color cameras
Xuan Dong 0001, Weixin Li 0001 |
Neurocomputing | 1 |
| 2018 | Differentiated Attentive Representation Learning for Sentence ClassificationabstractAttention-based models have shown to be effective in learning representations for sentence classification. They are typically equipped with multi-hop attention mechanism. However, existing multi-hop models still suffer from the problem of paying much attention to the most frequently noticed words, which might not be important to classify the current sentence. And there is a lack of explicitly effective way that helps the attention to be shifted out of a wrong part in the sentence. In this paper, we alleviate this problem by proposing a differentiated attentive learning model. It is composed of two branches of attention subnets and an example discriminator. An explicit signal with the loss information of the first attention subnet is passed on to the second one to drive them to learn different attentive preference. The example discriminator then selects the suitable attention subnet for sentence classification. Experimental results on real and synthetic datasets demonstrate the effectiveness of our model. Qianrong Zhou, Xiaojie Wang 0006, Xuan Dong 0001 |
IJCAI | 3 |
| 2018 | Object-Difference Attention: A Simple Relational Attention for Visual Question AnsweringabstractAttention mechanism has greatly promoted the development of Visual Question Answering (VQA). Attention distribution, which weights differently on objects (such as image regions or bounding boxes) in an image according to their importance for answering a question, plays a crucial role in attention mechanism. Most of the existing work focuses on fusing image features and text features to calculate the attention distribution without comparisons between different image objects. As a major property of attention, selectivity depends on comparisons between different objects. Comparisons provide more information for assigning attentions better. For achieving this, we propose an object-difference attention (ODA) which calculates the probability of attention by implementing difference operator between different image objects in an image under the guidance of questions in hand. Experimental results on three publicly available datasets show our ODA based VQA model achieves the state-of-the-art results. Furthermore, a general form of relational attention is proposed. Besides ODA, several other relational attentions are given. Experimental results show those relational attentions have strengths on different types of questions. Chenfei Wu, Jinlai Liu, Xiaojie Wang 0006, Xuan Dong 0001 |
ACM Multimedia | 4 |
| 2018 | Chain of Reasoning for Visual Question AnsweringabstractReasoning plays an essential role in Visual Question Answering (VQA). Multi-step and dynamic reasoning is often necessary for answering complex questions. For example, a question "What is placed next to the bus on the right of the picture?" talks about a compound object "bus on the right," which is generated by the relation . Furthermore, a new relation including this compound object is then required to infer the answer. However, previous methods support either one-step or static reasoning, without updating relations or generating compound objects. This paper proposes a novel reasoning model for addressing these problems. A chain of reasoning (CoR) is constructed for supporting multi-step and dynamic reasoning on changed relations and objects. In detail, iteratively, the relational reasoning operations form new relations between objects, and the object refining operations generate new compound objects from relations. We achieve new state-of-the-art results on four publicly available datasets. The visualization of the chain of reasoning illustrates the progress that the CoR generates new compound objects that lead to the answer of the question step by step. Chenfei Wu, Jinlai Liu, Xiaojie Wang 0006, Xuan Dong 0001 |
NeurIPS | 4 |
| 2018 | Ground-Truth Data Set and Baseline Evaluations for Base-Detail Separation Algorithms at the Part LevelabstractBase-detail separation is a fundamental image processing problem, which models the image by a smooth base layer for the coarse structure and a detail layer for the texturelike structures. Base-detail separation is hierarchical and can be performed from the fine level to the coarse level. The separation at coarse level, in particular at the part level, is important for many applications, but currently lacks ground-truth data sets that are needed for comparing algorithms quantitatively. Thus, we propose a procedure to construct such data sets and provide two examples: Pascal Part UCLA and Fashionista, containing 1000 and 250 images, respectively. Our assumption is that the base is piecewise smooth, and we label the appearance of each piece by a polynomial model. The pieces are objects and parts of objects obtained from human annotations. Finally, we propose a way to evaluate different separation methods with our data sets and compared the performances of seven state-of-the-art algorithms. Xuan Dong 0001, Boyan Bonev 0001, Weixin Li 0001, Weichao Qiu, Xianjie Chen, Alan L. Yuille |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2016 | Low lighting image enhancement using local maximum color value prior
Xuan Dong 0001, Jiangtao Wen |
Frontiers Comput. Sci. | 1 |
| 2015 | Region-based temporally consistent video post-processingabstractWe study the problem of temporally consistent video post-processing. Previous post-processing algorithms usually either fail to keep high fidelity or fail to keep temporal consistency of output videos. In this paper, we observe experimentally that many image/video enhancement algorithms enforce a spatially consistent prior on the enhancement. More precisely, within a local region, the enhancement is consistent, i.e., pixels with the same RGB values will get the same enhancement values. Using this prior, we segment each frame into several regions and temporally-spatially adjust the enhancement of regions of different frames, taking into account fidelity, temporal consistency and spatial consistency. User study, objective measurement and visual quality comparisons are conducted. The experimental results demonstrate that our output videos can keep high fidelity and temporal consistency at the same time. Xuan Dong 0001, Boyan Bonev 0001, Yu Zhu 0004, Alan L. Yuille |
CVPR | 1 |
| 2015 | Temporally consistent region-based video exposure correctionabstractWe analyze the problem of temporally consistent video exposure correction. Existing methods usually either fail to evaluate optimal exposure for every region or cannot get temporally consistent correction results. In addition, the contrast is often lost when the detail is not preserved properly during correction. In this paper, we use the block-based energy minimization to evaluate the temporally consistent exposure, which considers 1) the maximization of the visibility of all contents, 2) keeping the relative difference between neighboring regions, and 3) temporally consistent exposure of corresponding contents in different frames. Then, based on Weber contrast definition, we propose a contrast preserving exposure correction method. Experimental results show that our method enables better temporally consistent exposure evaluation and produces contrast preserving outputs. Xuan Dong 0001, Lu Yuan 0001, Weixin Li 0001, Alan L. Yuille |
ICME | 1 |
| 2015 | A pixel-based outlier-free motion estimation algorithm for scalable video quality enhancement
Xuan Dong 0001, Jiangtao Wen |
Frontiers Comput. Sci. | 1 |
| 2013 | Cross Segment Decoding for Improved Quality of Experience for Video ApplicationsabstractIn this paper, we present an improved algorithm for decoding live streamed or pre-encoded video bit streams with time-varying qualities. The algorithm extracts information available to the decoder from a high visual quality segment of the clip that has already been received and decoded, but was encoded independently from the current segment. The proposed decoder is capable of significantly improve the Quality of Experience of the user without incurring significant overhead to the storage and computational complexities of both the encoder and the decoder. We present simulation results using the HEVC reference encoder and standard test clips, and discuss areas of improvements to the algorithm and potential ways of incorporating the technique to a video streaming system or standards. Jiangtao Wen, Shunyao Li, Yao Lu 0006, Meiyuan Fang, Xuan Dong 0001, Huiwen Chang, Pin Tao |
DCC | 5 |
| 2011 | Fast efficient algorithm for enhancement of low lighting videoabstractWe describe a novel and effective video enhancement algorithm for low lighting video. The algorithm works by first inverting an input low-lighting video and then applying an optimized image de-haze algorithm on the inverted video. To facilitate faster computation, temporal correlations between subsequent frames are utilized to expedite the calculation of key algorithm parameters. Simulation results show excellent enhancement results and 4× speed up as compared with the frame-wise enhancement algorithms. Xuan Dong 0001, Yi Pang, Weixin Li 0001, Jiangtao Wen, Wei Meng 0001, Yao Lu 0006 |
ICME | 1 |