Submitted by Qi HU 2 SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces sssr-lab 3 3