Finetuning midddle layers (vision module layers = 22 to 26 +projector + language layers = 0 to 4)
HyperParameters of this model
lr: 2e-5
epochs: 3
batch size: 1
grad acumulation step: 1
mask = 262144 full vision module without patch embedding
Finetuning midddle layers (vision module layers = 22 to 26 +projector + language layers = 0 to 4)
HyperParameters of this model
lr: 2e-5
epochs: 3
batch size: 1
grad acumulation step: 1
mask = 262144 full vision module without patch embedding