LLM evaluation methodology, code security vulnerability detection, router hypothesis and structural prior injection, cross-domain generalization ceilings