Engage in multimedia chat with LLMs and ML models
Transcribe audio or YouTube video into text
Generate images from any text prompt