""" run_benchmark.py — Pre/Post training benchmark for the AgentOS agent. Runs N episodes per task against the local env (:7860), calls the local inference server (:8001) for agent actions, collects per-step reward data, and saves: - benchmark_results/results_