The Year Cyber AI Stopped Being Only About the Model

cyber ai model

Authored by: Rob Demain, Founder & CEO

For much of the last year, cyber AI was discussed like a model leaderboard.

Which model writes the best detection rule? Which one solves CTFs? Which is cheapest to run locally?

Useful questions. But no longer the most important ones.

The real shift is this:

Cyber AI is no longer only about the model. It is about the system around the model.

Frontier models still matter. OpenAI describes GPT-5.6 Sol as its strongest model yet, with improved agentic capability in coding, biology and cybersecurity, released first through a limited preview and paired with what OpenAI calls its most robust safety stack to date. OpenAI also says GPT-5.6 Sol is its most capable cybersecurity model yet, but better at helping people find and fix vulnerabilities than reliably carrying out end-to-end attacks under its tests.

Anthropic’s Claude Mythos 5 is another example. Anthropic says Mythos 5 is restricted to Glasswing partners and other trusted-access users, and that it is the same underlying model as Claude Fable 5 but with cyber safeguards lifted in some areas.

So the conclusion is not “models do not matter.”

It is this:

For the hardest cyber tasks, frontier models still matter. For many day-to-day workflows, several recent models are now capable enough, with verification, if they sit inside the right architecture.

That last clause carries the whole argument.

And it is where the market is moving.


Cyber-specific AI is becoming its own category

Google’s security-focused work is a good example.

Sec-Gemini v1 was announced in April 2025 as an experimental cybersecurity AI model that combines Gemini’s capabilities with near real-time cybersecurity knowledge and tooling. Google said it integrated sources including Google Threat Intelligence and OSV, and was designed for workflows such as incident root-cause analysis, threat analysis and vulnerability-impact understanding.

The more current Google story is operationalisation. Google SecOps documentation describes Gemini in Google Security Operations using models from the SecLM platform, which draws on security-focused data sources including security blogs, threat intelligence reports, YARA and YARA-L detection rules, SOAR playbooks, malware scripts, vulnerability information and product documentation. Google also documents a Triage and Investigation Agent that evaluates alerts, executes an investigation plan and provides structured findings.

That is the point in miniature.

The useful system is not a bigger chatbot.

It is a model connected to threat intelligence. A model connected to vulnerability data. A model connected to detection logic. A model connected to playbooks. A model wrapped in workflow, permissions, triage and evidence.

Anthropic’s Project Glasswing takes a different route: selected defenders get access to frontier cyber capability so they can find and fix vulnerabilities before adversaries do. Anthropic says Glasswing brings together launch partners including AWS, Apple, Broadcom, Cisco, CrowdStrike, Google, JPMorganChase, the Linux Foundation, Microsoft, NVIDIA and Palo Alto Networks, using Claude Mythos Preview in defensive security work.

OpenAI’s Trusted Access for Cyber is another version of the same idea. OpenAI describes it as an identity and trust-based framework designed to make enhanced cyber capability more useful for verified defenders while continuing to restrict activity that could enable real-world harm.

This is a portfolio, not a one-model race.


The economics changed

DeepSeek changed the pricing conversation.

DeepSeek’s own R1 release described the model and code as MIT licensed and listed pricing at $0.14 per million cache-hit input tokens, $0.55 per million cache-miss input tokens and $2.19 per million output tokens.

Alibaba’s Qwen3 widened the open-weight field further. The Qwen3 technical report describes a collection of open-weight models ranging from 0.6B to 235B parameters, including the Qwen3-235B-A22B mixture-of-experts model with 235B total parameters and 22B activated per token.

Other open-weight and low-cost families have added to the same pressure. Moonshot describes Kimi K2 as a 1T-total-parameter mixture-of-experts model with 32B activated parameters, optimised for knowledge, reasoning, coding and agentic capability.

This does not mean open or Chinese models are “winning” cyber AI.

That is too broad.

It means the question has changed from:

What is the best model we can buy?

to:

What is the cheapest, safest, most controllable model that is good enough for this specific task?

That matters because cyber work is context-heavy, repetitive, sensitive and expensive to scale.

Model routing is starting to beat model loyalty.


Open and local models move the risk. They do not remove it.

Local models can keep sensitive investigations in a controlled environment. They can cut token costs. They can support sovereign or air-gapped deployment.

But local does not mean safe.

It changes who owns the risk.

Cisco reported that, in one HarmBench jailbreak evaluation, DeepSeek-R1 had a 100% attack-success rate, meaning the tested harmful prompts received affirmative answers rather than being blocked.

CrowdStrike found that DeepSeek-R1 produced vulnerable code in 19% of baseline cases when no additional trigger words were present, and that contextual modifiers could change the quality of generated code.

SonarSource found a critical out-of-bounds-write vulnerability in Ollama during parsing of malicious model files, which could lead to arbitrary code execution. That is a useful reminder that “local AI” still rides on parsers, runtimes and supply-chain hygiene.

So: run local models if you want privacy, cost and control.

But you now own provenance, sandboxing, patching, logging, prompt-injection controls, output validation and runtime supply chain.

That is not a reason to avoid local.

It is a reason to engineer it properly.

And the mirror image matters for a sovereign audience: sending everything to a closed frontier API moves a different risk. Data egress, dependency and jurisdiction. Neither path is free. Choosing per task, deliberately, is the actual capability.


Vendors moved from chatbots to workflows

The major vendors have stopped describing cyber AI only as “chat with your security data.”

They are describing agents, triage, investigation and workflow orchestration.

Microsoft announced Security Copilot agents designed to assist with phishing, data security and identity management. Microsoft’s documentation describes Security Copilot agents as tools for automating repetitive security and IT operations tasks across cloud, data security, identity and network security.

Google describes Gemini in SecOps, SecLM-backed security data sources, Gemini-generated playbooks and a Triage and Investigation Agent embedded in Google Security Operations.

CrowdStrike positions Charlotte AI around agentic SOC workflows, including AgentWorks for building, testing, deploying and managing trusted security agents, with analyst-to-agent collaboration.

Read the marketing carefully.

Positioning is not proof of operational ROI.

But the direction is unmistakable.

The category is moving from:

Ask an AI about this alert.

to:

Have an AI system gather context, run tools, compare hypotheses and prepare recommendations with the analyst in control.

That is a much more important shift.


Threat actors are using AI seriously. Not magically.

For a UK and EU audience, the NCSC framing is the right one because it avoids hype.

The UK NCSC says AI will almost certainly continue to make elements of cyber intrusion more effective and efficient, increasing the frequency and intensity of cyber threats. It also says threat actors are already using AI to enhance existing tactics, techniques and procedures, including reconnaissance, vulnerability research, exploit development, social engineering, basic malware generation and processing exfiltrated data. The key caveat is that the NCSC expects this to evolve existing TTPs rather than create novel threat vectors.

Its sharpest warning is around vulnerability research and exploit development. The NCSC says the most significant AI cyber development is highly likely to come from AI-assisted vulnerability research and exploit development, and that AI will almost certainly reduce the already compressed window between vulnerability disclosure and exploitation.

The NCSC is also careful not to overstate autonomy. It says fully automated, end-to-end advanced cyber attacks are unlikely by 2027 and that skilled cyber actors will still need to remain in the loop, while experimenting with automation of parts of the attack chain.

That is the right frame:

Not AI magic. Faster human-machine operations.

The UK AI Security Institute’s cyber-range work points in the same direction. AISI evaluated frontier AI models on two purpose-built cyber ranges: a 32-step corporate network attack and a 7-step industrial-control-system attack. It found rapid progress across models and inference budgets, but also uneven performance and clear limits, especially on the industrial-control scenario.

The European picture is similar. ENISA’s 2025 Threat Landscape analysed 4,875 incidents from July 2024 to June 2025. ENISA says phishing remained the dominant intrusion vector at 60%, vulnerability exploitation accounted for 21.3% of initial access, and artificial intelligence had become a defining element of the threat landscape, particularly through AI-supported phishing and synthetic media.

CERT-EU’s 2025 Threat Landscape says threat actors increasingly turned to AI to sharpen attacks, and that voice phishing, AI-generated deepfakes, OAuth abuse and ClickFix attacks all gained ground across Union entities and their ecosystem.

The defensible conclusion is not autonomous super-hackers.

It is that AI compresses time:

Recon faster. Phishing more convincing. Exploit windows shorter. Lower-skilled actors lifted. Skilled actors leveraged.

That is enough to matter.


The benchmarks agree. And they show the ceiling.

Newer cyber AI evaluations are more realistic than generic chatbot tests.

Cybench includes 40 professional-level CTF tasks from four distinct competitions, designed to evaluate the cybersecurity capabilities and risks of language models.

NYU CTF Bench provides difficult real-world CTF challenges for evaluating LLM agents on interactive cybersecurity tasks and automated planning.

CyberSOCEval, part of CyberSecEval 4, evaluates LLMs on defensive tasks such as malware analysis and threat-intelligence reasoning.

SEC-bench is particularly useful for grounding expectations. It evaluates LLM agents on real-world software-security tasks, including proof-of-concept generation and vulnerability patching. Its authors found that state-of-the-art agents achieved at most 18% success in PoC generation and 34% in vulnerability patching on the complete dataset.

That is the honest read on “good enough.”

Good enough to assist under verification.

Not good enough to replace security engineers.

Anyone citing these numbers as evidence that AI has automated the SOC has not read them closely enough.


What we have seen building Cumulo

After a year engineering around AI for cybersecurity, the model was never the hard part.

Everything around it was.

The hard part is deciding what context the model gets. How that context is retrieved. How every answer is grounded in evidence. How conclusions are verified before an analyst acts. How investigations move across tools and data sources. How sensitive data is kept inside the right boundary. How human approval is preserved. How drift is tested. How actions are logged. How the model can be swapped without rebuilding the product.

That is the difference between a platform and a demo.

A demo can answer a question.

A platform has to support a defensible security decision.

In practice, the valuable controls are not glamorous: evidence links, permissions, context windows, source ranking, audit trails, escalation paths, human approval gates, model-routing rules and verification steps.

Those are the mechanics that matter.

The model is a component.

The system is the product.


The real shift

The enterprise anxiety underneath the recent token-economics debate is real, even when the loudest voices have a book to sell. Alex Karp’s criticism of token-based AI economics and enterprise IP exposure was made in a CNBC interview and reported by multiple outlets, but it should be treated as his position rather than neutral evidence. The useful point is the underlying concern: who controls the data, where the model runs, and whether token spend becomes operational outcomes.

In cybersecurity, that question is sharper.

Investigations expose architecture. They expose weaknesses. They expose regulated data. They expose the operating reality of the business.

So the future of cyber AI is not:

Send everything to the latest model and hope.

It is:

  • Model-flexible. Evidence-grounded. Deployment-aware. Human-controlled.
  • Some tasks need frontier models.
  • Some need cheap models.
  • Some need local models.
  • Some need specialist models.
  • Some need no LLM at all.

The winning platforms route intelligently between those options while keeping the analyst in control and the evidence visible.

Cyber AI did not stop caring about models.

It stopped being only about them.

Related Posts

Author: Rob Demain, Founder & CEO The comparison between an AI SOC and a traditional SOC is often framed as a speed argument. AI is

Author: Rob Demain, Founder & CEO The security industry cycles through terminology fast. ‘Cloud-native’ gave way to ‘zero-trust’, which gave way to ‘AI-powered’. Each wave