Finetuning all layers except ( initial_patch_embedding and position embedding: first 3 layers) HyperParameters of this model lr: 2e-5 epochs: 24 batch size: 1 grad acumulation step: 4