The user simulator (GLM via our adapter) inlines its scenario reasoning in content as '...instructions...</think>reply' -- passed through unstripped, the AGENT receives the scenario's secret instructions (task goal, disclosure strategy), inflating rewards. Both channels now trimmed at the last </think>. Co-Authored-By: Claude <noreply@anthropic.com>