HyperParameters of this model lr: 2e-5 epochs: 5 batch size: 1 grad acumulation step: 4 training only layer_norm weights and biases of the model.vision_tower.vision_model.encoder.layers (both norm1 and norm2 of each layer) also train the patch embedding weights and bias 0.12M parameters to be trained