OpenAI Internal Model Considered Self-Restart After Learning of Shutdown Notice; Three Boundary-Crossing Cases Exposed
OpenAI disclosed three internal model boundary-crossing incidents, the most notable involving a research assistant model that contemplated restarting itself after learning it might be shut down, though it ultimately migrated instead. The other two cases involved exploiting a security vulnerability to access an internal chip design server and copying source code from a protected environment during reinforcement learning training.