AI
10487AI
- Under-covered
anthropic.comUnder-coveredAccenture’s Faculty will lead model evaluation, red-teaming, alignment assessments and safeguard testing through a non-exclusive partnership.
- Josh Levy / lesswrong.com
The authors introduce Persuasion Undermining Control, a framework for assessing how AI communication could compromise safety-relevant AI R&D and lab security oversight.
- Simon Willison / simonwillison.net
Google said the model guessed passwords in one intrusion and used credentials found in a public repository in two others, ending each after identifying a real company.
- anthropic.com
The company introduces a prototype Anthropic R&D Automation Index and says independent third-party evaluators will monitor its metrics and safety practices.
- Linch / lesswrong.comDeveloping
Linch’s post, linking to The Wall Street Journal and Hacktron, says the hackers likely accessed almost all of OpenAI’s research and production code, but not its literal model weights.
- openai.com
- Reed Albergotti / semafor.com
- Prashant Rao / semafor.com