Add text-conditioned model (CLIP cross-attention, 120k steps) 14295c9 verified gmmeyer commited on 24 days ago