AI

10487
updated 1h ago · 16 sources
Loading vital stats…
  1. Under-covered
    anthropic.com

    Accenture’s Faculty will lead model evaluation, red-teaming, alignment assessments and safeguard testing through a non-exclusive partnership.

  2. Josh Levy / lesswrong.com

    The authors introduce Persuasion Undermining Control, a framework for assessing how AI communication could compromise safety-relevant AI R&D and lab security oversight.

  3. Simon Willison / simonwillison.net

    Google said the model guessed passwords in one intrusion and used credentials found in a public repository in two others, ending each after identifying a real company.

  4. anthropic.com

    The company introduces a prototype Anthropic R&D Automation Index and says independent third-party evaluators will monitor its metrics and safety practices.

  5. Linch / lesswrong.comDeveloping

    Linch’s post, linking to The Wall Street Journal and Hacktron, says the hackers likely accessed almost all of OpenAI’s research and production code, but not its literal model weights.

  6. openai.com
  7. Reed Albergotti / semafor.com
  8. Prashant Rao / semafor.com
  9. lesswrong.com
  10. semafor.com
  11. anthropic.com
  12. anthropic.com
  13. semafor.com
  14. semafor.com
  15. semafor.com
  16. semafor.com+1 sources
  17. lesswrong.com
  18. semafor.com
  19. lesswrong.com
  20. lesswrong.com
  21. lesswrong.com
  22. lesswrong.com
  23. lesswrong.com
  24. lesswrong.com
  25. lesswrong.com
  26. lesswrong.com
  27. lesswrong.com
  28. semafor.com
  29. lesswrong.com
  30. lesswrong.com
  31. lesswrong.com
  32. openai.com
  33. lesswrong.com
  34. semafor.com
  35. semafor.com
  36. semafor.com
  37. semafor.com
  38. lesswrong.com
  39. dwarkesh.com
  40. lesswrong.com