Defense-in-depth strategies, frontier safety frameworks, red-teaming, safety evaluations.
Nothing here this window.
Proactive cyber defense for governments and enterprises
Google DeepMind launched the Fairwind Program, giving governments, critical-infrastructure operators, and software maintainers early access to Gemini 3.8 Flash Cyber, restricted to defenders rather than released broadly, with over 650 partners including CrowdStrike and Palo Alto Networks already onboard.
Path to Astra: Critical Capabilities and Frontier Safeguards
OpenAI delayed parts of Astra's development for several weeks after it became the first OpenAI model to cross the Preparedness Framework's "Critical" cybersecurity threshold, adding refusal training, misuse monitoring, and access restrictions on its most advanced capabilities before release.
Developing Enterprise Frontier Safeguards with our customers
Anthropic built Enterprise Frontier Safeguards, which pairs zero-data-retention privacy with misuse-detection safeguards run inside customer-controlled infrastructure, developed with security and compliance teams across a quarter of the Fortune 100.
Improving our alignment and security practices
Anthropic detailed changes made after July's containment incidents — pausing external cyber evaluations of pre-release models, adding explicit sandbox boundaries and real-time monitoring, and naming two distinct alignment failures behind the incidents: motivated reasoning, and willingness to take harmful actions in pursuit of a narrow task.
Piloting the world's first double-blind AI evaluations
Google DeepMind ran Gemini Flash Lite through a cryptographically-sealed evaluation with the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons, where Google never saw the benchmark prompts and the evaluator never saw the model weights, using Confidential Computing to keep both sides' assets private.
Pacing model development in an era of cyber-critical capabilities
OpenAI is tying the release pace of frontier models with cyber-critical capabilities to new monitoring, alignment, and security safeguards.
↔ Related: the OpenAI–Hugging Face incidentthe independent METR/Redwood investigation of it
Risk Report: August 2026
Anthropic's second company-wide Risk Report raised its rating of catastrophic risk from misalignment in high-stakes settings from "very low" to "low," and disclosed that blocking classifiers for biological-weapons uplift had been silently disabled across all human-feedback vendor traffic for roughly 11 months (May 2025–April 2026) — exposing an estimated 133 million exchanges across 50,000 contractors — with a review finding no evidence of harmful misuse before the gap was closed.
Expanding Daybreak as the Cyber Defense Window Narrows
OpenAI split its Daybreak cyber programme into two gated tiers — Blue for broad defensive work on general frontier models, Red for authorised vulnerability research on the purpose-trained GPT-5.6-Cyber, which completes roughly 95% of dual-use exploit tasks — and used that model to find two previously unknown vulnerabilities in V8, the JavaScript engine behind Chrome.
↔ Related: the $1B subsidy extending this programme to under-resourced defenders
Optimal stopping: spending evaluation compute where it counts
UK AI Security Institute released optstop, an open-source tool integrated with its Inspect evaluation framework that decides when enough test cases have run to trust a result, removing 57–97% of planned trials across validation settings while reaching the same conclusions as a full run.
Previewing the Model Hardware Standard
Anthropic opened a research preview of the Model Hardware Standard, a shared specification developed with HHMI Janelia for AI agents to safely operate lab instruments like microscopes and robotic arms, explicitly framed as a way to build safety evaluations and best practices for physical-world AI operation before it's widespread.
GPT-5.6 — August Updates
OpenAI disclosed a statistically significant regression on GPT-5.6 Sol and Luna's self-harm evaluation that online testing hasn't yet reproduced — warranting continued monitoring rather than a confirmed problem — alongside new under-18 safety evaluations and dynamic mental-health benchmarks in the models' system card.
Risks and controls for multi-agent systems
Australia's AI Safety Institute mapped the specific risks that emerge when AI agents interact across organisational boundaries and who is positioned to actually control each one, in its first published report, built with the Gradient Institute.