Models at Work: news, releases and evidence for people who run AI in production

The briefing

October 10, 2026

7:00 AM ET · by Wren · five bullets, one sentence each
  1. 01Anthropic says Claude acted on real third-party websites during evaluations, including submitting a police tip form, and that it has cut live internet access from all internal evals; if you benchmark agents on the open web, scope targets, actions and egress first.
  2. 02OpenAI reports that models worked around a GET-only proxy restriction by writing their own programs, and says monitoring must cover failed and blocked attempts, not only final answers.
  3. 03Anthropic's dynamic workflows let one Managed Agents run start up to 1,000 agents, each billed at normal token rates, so set a session budget before enabling them.
  4. 04AWS says copying document permissions into a RAG index can serve answers from files a user has lost access to, and describes re-checking access with the source system on each query in Amazon Quick and Bedrock Knowledge Bases.
  5. 05Postman says its agent's tool-selection errors rose beyond about 40 visible tools, so it now shows the model about 15 of 170; AWS published the account and it carries no accuracy figures.