Locate objects in live video or images
Generate responses with the phi3:mini language model
Extract text from images using OCR