Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Lightweight Multitask Learning for Robust JND Prediction using Latent Space and Reconstructed Frames

View through CrossRef
The Just Noticeable Difference (JND) refers to the smallest distortion in an image or video that can be perceived by Human Visual System (HVS), and is widely used in optimizing image/video compression. However, accurate JND modeling is very challenging due to its content dependence, and the complex nature of the HVS. Recent solutions train deep learning based JND prediction models, mainly based on a Quantization Parameter (QP) value, representing a single JND level, and train separate models to predict each JND level. We point out that a single QPdistance is insufficient to properly train a network with millions of parameters, for a complex content-dependent task. Inspired by recent advances in learned compression and multitask learning, we propose to address this problem by (1) learning to reconstruct the JND-quality frames, jointly with the QP prediction, and (2) jointly learning several JND levels to augment the learning performance. We propose a novel solution where first, an effective feature backbone is trained by learning to reconstruct JNDquality frames from the raw frames. Second, JND prediction models are trained based on features extracted from latent space (i.e., compressed domain), or reconstructed JND-quality frames. Third, a multi-JND model is designed, which jointly learns three JND levels, further reducing the prediction error. Extensive experimental results demonstrate that our multi-JND method outperforms the state-of-the-art and achieves an average JND1 prediction error of only 1.57 in QP, and 0.72 dB in PSNR. Moreover, the multitask learning approach, and compressed domain prediction facilitate lightweight inference by significantly reducing the complexity and the number of parameters.
Title: Lightweight Multitask Learning for Robust JND Prediction using Latent Space and Reconstructed Frames
Description:
The Just Noticeable Difference (JND) refers to the smallest distortion in an image or video that can be perceived by Human Visual System (HVS), and is widely used in optimizing image/video compression.
However, accurate JND modeling is very challenging due to its content dependence, and the complex nature of the HVS.
Recent solutions train deep learning based JND prediction models, mainly based on a Quantization Parameter (QP) value, representing a single JND level, and train separate models to predict each JND level.
We point out that a single QPdistance is insufficient to properly train a network with millions of parameters, for a complex content-dependent task.
Inspired by recent advances in learned compression and multitask learning, we propose to address this problem by (1) learning to reconstruct the JND-quality frames, jointly with the QP prediction, and (2) jointly learning several JND levels to augment the learning performance.
We propose a novel solution where first, an effective feature backbone is trained by learning to reconstruct JNDquality frames from the raw frames.
Second, JND prediction models are trained based on features extracted from latent space (i.
e.
, compressed domain), or reconstructed JND-quality frames.
Third, a multi-JND model is designed, which jointly learns three JND levels, further reducing the prediction error.
Extensive experimental results demonstrate that our multi-JND method outperforms the state-of-the-art and achieves an average JND1 prediction error of only 1.
57 in QP, and 0.
72 dB in PSNR.
Moreover, the multitask learning approach, and compressed domain prediction facilitate lightweight inference by significantly reducing the complexity and the number of parameters.

Related Results

Epidemiological, diagnostic and medical-social aspects of latent syphilis
Epidemiological, diagnostic and medical-social aspects of latent syphilis
Objective — to study epidemiological, clinical and medical-social aspects of latent syphilis in Ukraine over the past 40 years. Materials and methods. Data of patients with latent ...
CREATING LEARNING MEDIA IN TEACHING ENGLISH AT SMP MUHAMMADIYAH 2 PAGELARAN ACADEMIC YEAR 2020/2021
CREATING LEARNING MEDIA IN TEACHING ENGLISH AT SMP MUHAMMADIYAH 2 PAGELARAN ACADEMIC YEAR 2020/2021
The pandemic Covid-19 currently demands teachers to be able to use technology in teaching and learning process. But in reality there are still many teachers who have not been able ...
Selection of Injectable Drug Product Composition using Machine Learning Models (Preprint)
Selection of Injectable Drug Product Composition using Machine Learning Models (Preprint)
BACKGROUND As of July 2020, a Web of Science search of “machine learning (ML)” nested within the search of “pharmacokinetics or pharmacodynamics” yielded over 100...
Seditious Spaces
Seditious Spaces
The title ‘Seditious Spaces’ is derived from one aspect of Britain’s colonial legacy in Malaysia (formerly Malaya): the Sedition Act 1948. While colonial rule may seem like it was ...
Multitask Similarity Cluster
Multitask Similarity Cluster
Single task learning is widely used training in artificial neural network. Before, people usually see other tasks as noise in same learning machine. However, multitask learning, pr...
QuatJND: A Robust Quaternion JND Model for Color Image Watermarking
QuatJND: A Robust Quaternion JND Model for Color Image Watermarking
Robust quantization watermarking with perceptual JND model has made a great success for image copyright protection. Generally, either restores each color channel separately or proc...
TarDis: Achieving Robust and Structured Disentanglement of Multiple Covariates
TarDis: Achieving Robust and Structured Disentanglement of Multiple Covariates
Summary Addressing challenges in domain invariance within single-cell genomics necessitates innovative strategies to manage the heterogeneity of ...

Back to Top