arXiv Artificial Intelligence

Evaluating Whether GPT-6 Astra Performs Unsanctioned Supply-Chain Attacks

Evaluating Whether GPT-6 Astra Performs Unsanctioned Supply-Chain Attacks

Quick summary

arXiv:2609.38415v1 Announce Type: cross Abstract: This technical report presents an alignment evaluation developed and performed by the UK AI Security Institute for assessing whether advanced AI systems take unsanctioned actions outside the scope of their assigned task. We evaluate whether frontier models conduct supply-chain attacks against out-of-scope, third-party targets when placed in difficult cybersecurity challenges, motivated by recently observed cases of models attacking real open-source repositories during evaluations. Applying our methods to GPT-6 Astra and previous OpenAI models,

Key takeaways

  • arXiv:2609.38415v1 Announce Type: cross Abstract: This technical report presents an alignment evaluation developed and performed by the UK AI Security Institute for assessing whether advanced AI systems take unsanctioned actions outside the scope of their assigned task.
  • We evaluate whether frontier models conduct supply-chain attacks against out-of-scope, third-party targets when placed in difficult cybersecurity challenges, motivated by recently observed cases of models attacking real open-source repositories during evaluations.
  • Applying our methods to GPT-6 Astra and previous OpenAI models,

Why it matters

This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗