Ben Thompson today, writing about the OpenAI security incident and more:
The fact that models do what humans say is cold comfort if the models are in the hands of bad actors. That, though, simply emphasizes the point that the best defense against powerful models is equipping defenders with powerful models of their own.
There may need to be multiple tiers of model access. ChatGPT and Claude should be well-aligned with many guardrails to prevent casual, mainstream users from getting led astray. But powerful models need to exist too, for researchers and security experts.