train only the projector fn (last layernorm of vision + projector layers + embed tokens of language model) : 4 layers HyperParameters of this model lr: 2e-5 epochs: 10 batch size: 1 grad acumulation step: 4