Intent detection, agent safety, behavioral attestation, small language models, fine-tuning, model evaluation