Switch inference to in-process ZeroGPU (inference.py); drop Modal calls 4fc2e3f verified trainig commited on Jun 15