ViT-B/16 fine-tuned on Stanford 40 Actions

Vision Transformer ViT-B/16 fine-tuned on the Stanford 40 Actions dataset for action recognition in still images.
This model reaches 91.5% accuracy on the held-out test split.


Model description

  • Architecture: ViT-Base, patch size 16 (224ร—224 resolution)
  • Base model: google/vit-base-patch16-224-in21k
  • Task: Single-label image classification (40 human actions)
  • Input: RGB image with a person performing an action
  • Output: One of 40 action labels (e.g. applauding, rowing a boat, texting message)

How to use

from PIL import Image
import requests
from transformers import AutoImageProcessor, AutoModelForImageClassification

model_id = "lucasddmc/vit-b16-stanford40-actions"

processor = AutoImageProcessor.from_pretrained(model_id)
model = AutoModelForImageClassification.from_pretrained(model_id)

url = "https://example.com/some_action_image.jpg"
image = Image.open(requests.get(url, stream=True).raw).convert("RGB")

inputs = processor(images=image, return_tensors="pt")
outputs = model(**inputs)
logits = outputs.logits
predicted_class_idx = logits.argmax(-1).item()

print(model.config.id2label[predicted_class_idx])

Training details (summary)

  • Dataset: Stanford 40 Actions
  • Framework: Trained with PyTorch + timm; exported to Hugging Face Transformers format (config.json/model.safetensors) for inference
  • Preprocessing / Augmentation (typical for ViT):
    • Resize to 256ร—256
    • RandomResizedCrop to 224ร—224
    • RandomHorizontalFlip
    • Normalization with ImageNet mean/std
  • Loss: CrossEntropyLoss (with/without label smoothing, depending on your setup)
  • Optimizer: AdamW
    • Body LR: 1e-5โ€“3e-5
    • Head LR: 1e-4โ€“3e-4
    • Weight decay: 0.01โ€“0.05
  • Scheduler: CosineAnnealingLR
  • Epochs: ~20โ€“50 (adjust based on your actual training)
  • Hardware: 1x GPU (specify if desired: V100, A100, etc.)

Update these values with the exact numbers from your experiment.


Evaluation

  • Metric: Top-1 accuracy
  • Split: Stanford 40 Actions test
  • Result: 91.5%
Split Metric Value
Test Accuracy 0.915

Intended uses & limitations

Intended uses

  • Research in action recognition on still images
  • Baseline / feature extractor for computer vision experiments
  • Educational use (example of ViT fine-tuning on a human actions dataset)

Limitations

  • Trained only on Stanford 40 Actions (limited context/cultural diversity).
  • Assumes a single primary action per image.
  • Not recommended for use in sensitive scenarios (surveillance, critical decision-making, etc.) without careful evaluation.

Citation

If you use this model, please cite the dataset:

@INPROCEEDINGS{6126386,
  author={Yao, Bangpeng and Jiang, Xiaoye and Khosla, Aditya and Lin, Andy Lai and Guibas, Leonidas and Fei-Fei, Li},
  booktitle={2011 International Conference on Computer Vision}, 
  title={Human action recognition by learning bases of action attributes and parts}, 
  year={2011},
  volume={},
  number={},
  pages={1331-1338},
  keywords={Humans;Detectors;Image reconstruction;Vectors;Feature extraction;Noise;Image recognition},
  doi={10.1109/ICCV.2011.6126386}
}

And this model's page on Hugging Face:

@misc{vit_b16_stanford40_actions,
  title={ViT-B/16 fine-tuned on Stanford 40 Actions},
  author={Lucas Dantas de Moura Carvalho},
  year={2025},
  howpublished={\url{https://huggingface.co/lucasddmc/vit-b16-stanford40-actions}}
}
Downloads last month
8
Safetensors
Model size
85.8M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for lucasddmc/vit-b16-stanford40-actions

Finetuned
(2583)
this model

Evaluation results