Synthetic user–agent trajectories with tool calls, tool results, and conduct grades for post-training and evaluation