Better image classification in new domain by focusing on foreground with RobustViT

Better image classification in new domain by focusing on foreground with RobustViT

Optimizing Relevance Maps of Vision Transformers Improves Robustness
arXiv paper abstract https://arxiv.org/abs/2206.01161
arXiv PDF paper https://arxiv.org/pdf/2206.01161.pdf
GitHub https://github.com/hila-chefer/RobustViT
Online demo https://huggingface.co/spaces/Hila/RobustViT

… visual classification models often rely mostly on the image background, neglecting the foreground, which hurts their robustness to distribution changes.

… propose to monitor the model’s relevancy signal and manipulate it such that the model is focused on the foreground object.

This is done as a finetuning step, involving relatively few samples consisting of pairs of images and their associated foreground masks.

… encourage the model’s relevancy map (i) to assign lower relevance to background regions, (ii) to consider as much information as possible from the foreground, and (iii) … encourage the decisions to have high confidence.

When applied to Vision Transformer (ViT) models, a marked improvement in robustness to domain shifts is observed.

… foreground masks can be obtained automatically, from a self-supervised variant of the ViT model itself; therefore no additional supervision is required.

Stay up to date. Subscribe to my posts https://morrislee1234.wixsite.com/website/contact
Web site with my other posts by category https://morrislee1234.wixsite.com/website

LinkedIn https://www.linkedin.com/in/morris-lee-47877b7b

Photo by Jacinto Diego on Unsplash

--

--

A computer vision consultant in artificial intelligence and related hitech technologies 37+ years. Am innovator with 66+ patents and ready to help a firm's R&D.

Get the Medium app

A button that says 'Download on the App Store', and if clicked it will lead you to the iOS App store
A button that says 'Get it on, Google Play', and if clicked it will lead you to the Google Play store
AI News Clips by Morris Lee: News to help your R&D

A computer vision consultant in artificial intelligence and related hitech technologies 37+ years. Am innovator with 66+ patents and ready to help a firm's R&D.