Skip to main content
GIS-Schools GIS-Schools
Classification ยท GeoAI

ViT Base Patch16 - ImageNet Image Classification

A patch-based Transformer counterpart to ResNet. Explore whole-image classification, transfer learning and the difference between image labels and mapped land cover.

Classification task illustration
Family: Vision TransformerArchitecture: 16-pixel patches + Transformer encoder + class tokenVersion: Base patch16 224 / HF 3f49326License: Apache-2.0 (provider model card)
ViT Base Patch16 - ImageNet Image Classification

Exact checkpoint

google/vit-base-patch16-224 at revision 3f49326eb077187dfe1c2a2bb15fbd74e6ab91e3. This checkpoint was pretrained on ImageNet-21k and fine-tuned on ImageNet-1k. The provider card credits conversion through timm from JAX into PyTorch; it is not a claim that the original authors directly released this safetensors serialization.

How it works

The image becomes a sequence of square patch embeddings. Positional information retains their arrangement, while attention allows distant patches to exchange information. A classification token summarizes the scene for the prediction head.

Typical input and output

This selected variant uses RGB 224 by 224 input and 16 by 16 patches. Use its image processor. Output is an ImageNet class-score vector for the whole image, not a spatial label grid.

Strengths

Attention provides a different way to combine context across a scene. Keeping a Transformer and a CNN baseline helps separate architecture effects from dataset or split effects.

GIS and remote-sensing use

Fine-tune for scene classification of spatially defined chips. Decide whether the desired label represents dominant cover, land use or presence of a feature before annotating. A scene classifier cannot replace a segmentation model when individual pixels need labels.

Limitations

The provided classes describe ImageNet images rather than geospatial land-use categories. Small objects can be diluted when a large image is resized. Attention visualizations are not automatically reliable explanations or GIS boundaries.

Training and fine-tuning

Replace the classifier with the target label set. Use consistent chip extent and processor settings, then compare frozen-backbone and fine-tuned runs. Keep neighboring or overlapping chips in the same split. If experimenting with larger image sizes, follow the implementation guidance for positional embeddings instead of silently changing resolution.

Framework and hardware

PyTorch and Hugging Face Transformers; original JAX implementation is also linked. Load the exact revision and its processor. A GPU is useful for fine-tuning; practical batch size depends on resolution and device memory. No local inference speed was measured.

Files, license and download

Official external checkpoint: model.safetensors is 346,293,852 bytes, above the existing 100 MiB resource ceiling. The provider declares Apache-2.0; size is the reason this checkpoint is not mirrored here.

Before using a result

  1. Keep imagery rights, acquisition dates and preprocessing with each experiment.
  2. Separate locations before tiling to avoid train/test leakage.
  3. Inspect failures and uncertain outputs alongside the original imagery.
  4. Preserve coordinate transforms and verify units before exporting a GIS layer.

Validation status: Source provenance and website delivery are checked. GISSchools has not run model inference, training or an accuracy benchmark for this record. Feature images and gallery diagrams are explanatory illustrations, not predictions.

Official references

Examples and images

Downloads and resources

External resourceOfficial Model WeightsProvider: Google model namespace / Hugging Face conversionVersion: Base patch16 224 / HF 3f49326License: Apache-2.0 (provider model card)

Official external checkpoint: model.safetensors is 346,293,852 bytes, above the existing 100 MiB resource ceiling. The provider declares Apache-2.0; size is the reason this checkpoint is not mirrored here.

Open resource ↗
External resourceOfficial RepositoryProvider: Google model namespace / Hugging Face conversion
Open resource ↗
External resourceDocumentationProvider: Google model namespace / Hugging Face conversion
Open resource ↗
External resourceResearch PaperProvider: Google model namespace / Hugging Face conversion
Open resource ↗
External resourceLicense and Usage TermsProvider: Google model namespace / Hugging Face conversionLicense: Apache-2.0 (provider model card)
Open resource ↗
Discussion

Comments

0

No comments yet. Start the discussion.

Join the discussion

Leave a reply

Your email address will not be published. Required fields are marked.