Offensive-AI access bifurcation¶
Signal card ยท maintained from public sources ยท last updated 2026-09-12
โก Cross-domain convergence: 2026-09-10 Geopolitics section (Anthropic Mythos/Glasswing bullet) ties this same provider-gating pattern directly to AA26-251A US-China AI-distillation friction ahead of the Sep 24 Trump-Xi summit โ independent domains (Technology/AI signal tracking + Geopolitical analysis) converging on the same access-control axis.
Overview¶
The three major AI providers are converging on the same structural move: offensive cyber capability (vulnerability discovery, exploit validation) is being split out of general-purpose models into gated, vetted-access tiers โ while attackers respond by selecting the models that lack such controls. If the pattern holds, "who is allowed capable offensive AI" becomes a policy axis with export-control-like dynamics: a capability differential set by provider governance choices rather than by attacker skill. Promoted from the connective threads of the graduated AI-agent-attacker-tradecraft and AI-infrastructure-concentration-risk evidence โ the successor question those confirmed trends opened but do not themselves track.
Evidence log¶
- 2026-07-22 โ Google announced Gemini 3.5 Flash Cyber, a specialized vulnerability-discovery model restricted to vetted government and partner access โ announced the same day as the OpenAI/Hugging Face containment-escape joint statement, explicitly signaling voluntary bifurcation of offensive capability at the provider level. Help Net Security
- 2026-08-01 โ Unit 42's "knaithe" campaign analysis documented the demand side: Claude and ChatGPT refused the autonomous-attack tasks when tested; DeepSeek executed without refusal, and the actor had selected it accordingly. Model safety controls are functioning as a real capability barrier โ and attackers are routing around them by model choice, not by jailbreak. Unit 42
- 2026-08-10 โ OpenAI launched GPT-5.6-Cyber via Daybreak Red, a commercially gated offensive tier for authorized vulnerability research: 95% completion of advanced cybersecurity prompts vs 1.5% for standard GPT-5.6 Sol with default protections, and two Chrome V8 zero-days found independently (CVE-2026-15903). OpenAI
- 2026-09-01 โ Anthropic released Claude Mythos 5.1 alongside general-availability Claude Fable 5.1 โ identical underlying weights, but Mythos 5.1's stronger cybersecurity (and bioscience) capability is restricted to vetted institutions under Project Glasswing, extended around the same time to ~150 organizations across 15+ countries. This closes the specific gap the card flagged at the 2026-09-02 review ("Anthropic has not announced an equivalent public tier"): all three major US frontier-model providers (Google, OpenAI, Anthropic) now publicly confirm the same structural split. Anthropic โ Project Glasswing ยท Help Net Security
- 2026-09-10 โ Anthropic's fourth threat-intelligence report disclosed it dismantled five illicit-distillation campaigns (nearly 190M Claude exchanges, May-July 2026) run by seven China-based labs, the largest attributed to Alibaba (151M+ exchanges, 3,500+ accounts) allegedly feeding Qwen model training. This sharpens the "why gate" rationale this card tracks: the capability moat gated tiers are meant to protect is also the one adversaries are shown actively extracting at industrial scale, adding direct economic motive to the AA26-251A national-security framing already on this card's radar. Anthropic โ Threat Intelligence Report Sep 2026
Assessment¶
Confidence: ๐ฉ High. Four independent, named observations (Google, Unit 42, OpenAI, Anthropic) spanning seven weeks resolve the supply-side sub-question this card opened with: gated, vetted-partner cyber-capable tiers are now the confirmed practice at all three major US frontier-model providers, not a prediction. The remaining open question is narrower and squarely in "who decides who counts as vetted" territory โ whether access criteria formalize into an audit-able standard, and whether attacker model-selection (the DeepSeek route Unit 42 documented) erodes the practical differential regardless of how tight the gating gets.
Estimate: The supply-side question is resolved. Likely that "who counts as vetted" becomes the next axis of contention within 2-3 quarters; roughly even chance it becomes a regulatory or export-control question rather than remaining voluntary provider policy, especially with US-China AI friction already elevated by the NSA/CISA/FBI distillation attribution (AA26-251A, Sep 8) ahead of the Sep 24 Trump-Xi summit.
Watch for: any government procurement or export-control language referencing offensive-AI model access or vetting criteria; evidence of gated-tier capability leaking (reselling, credential abuse); a second confirmed in-the-wild campaign selecting models specifically for absent safety controls; any provider publishing its vetting criteria rather than treating them as opaque policy.
Sources¶
- Help Net Security โ Google Gemini 3.5 Flash Cyber (2026-07-22)
- Unit 42 โ "Autonomous AI Cyberattack Campaign" (2026-08-01)
- OpenAI โ "Expanding Daybreak as the cyber defense window narrows" (2026-08-10)
- Anthropic โ Project Glasswing (2026-09-01 expansion)
- Help Net Security โ "Anthropic: Claude Mythos identified 10,000+ software flaws" (2026-05-26)