LLM inference, OpenAI-compatible APIs, DeepSeek models, tool calling, structured outputs, and scalable GPU compute