
Aikido Security says it used 11.7 billion tokens in an effort to identify the strongest cyber AI model, putting the cost and complexity of model evaluation at the center of a cybersecurity research story.
The report, titled “We burned 11.7bn tokens to find the best cyber AI model,” is the only substantive item in the supplied source cluster. Its full article text is unavailable, so the specific models tested, tasks used, spending, scoring method, and final recommendation cannot be independently established from the available evidence. The two indexed source entries are also duplicates, not separate confirmations.
For AI builders and security teams, the headline still points to a practical problem: selecting a model for security work is not the same as choosing a model from a general-purpose leaderboard. Results can depend on the kind of vulnerability analysis, codebase, alert data, tool access, context window, and human review used during testing.
The confirmed fact is narrow. Aikido Security published an account describing an 11.7-billion-token evaluation intended to find the best cyber AI model. The available evidence does not identify whether that figure represents input tokens, output tokens, or a combined total. It also does not disclose whether the work involved hosted application programming interfaces, locally deployed models, agent loops, or a mixture of approaches.
That missing detail matters because token consumption is not a direct measure of research quality. A long evaluation may reflect broad test coverage, repeated prompting, automated retries, large code repositories, or inefficient workflows. Conversely, a smaller test can miss important failure modes. Without the experimental protocol, readers cannot determine whether the reported effort was a controlled benchmark, an internal product exercise, or a broader trial of models in operational security workflows.
The source should therefore be treated as a vendor-reported evaluation claim rather than an independent industry benchmark. Nothing in the supplied material confirms that Aikido Security’s eventual ranking applies across security teams, programming languages, threat categories, or deployment environments.
Token usage has become a material part of AI product economics. In a cyber defense workflow, a model may process source code, dependency files, vulnerability descriptions, logs, tickets, remediation suggestions, and tool results. An agent that repeatedly revisits those materials can consume far more tokens than a single prompt-and-answer interaction.
That makes the reported figure relevant even without a disclosed winner. It suggests that serious model selection can require a substantial evaluation budget when teams test multiple systems across realistic tasks. The cost is not limited to model access. Engineering teams also need test harnesses, representative data, scoring rules, sandboxed tools, and reviewers who can distinguish a plausible answer from a safe and correct one.
For enterprise AI buyers, the key question is not simply which model achieved the highest score. It is whether the model can deliver reliable results at an acceptable cost while preserving data controls. A model that performs well on a narrow vulnerability task may be unsuitable if it requires sending sensitive source code to an external provider, produces difficult-to-audit recommendations, or fails when a repository exceeds its usable context.
A useful cyber AI model evaluation should separate several capabilities. Code understanding and vulnerability discovery are different from explaining exploitability, proposing a patch, validating that patch, or prioritizing findings for a security team. Tool use introduces another layer: an agent may locate more issues when it can search files, run tests, inspect dependencies, or query a scanner, but those permissions also create safety and governance concerns.
The available source does not say which of these dimensions Aikido Security measured. That prevents a meaningful comparison between the reported result and existing AI benchmarks. It also leaves open whether the experiment measured accuracy, recall, false-positive rates, remediation quality, latency, cost, or human productivity.
This distinction is important for product teams building coding assistants and AI security tools. A model can appear strong in a demonstration while creating operational friction through excessive alerts, unsafe code changes, or recommendations that require extensive manual review. Security workflows generally value consistent, explainable decisions over fluent responses alone.
Builders evaluating AI models for security should view Aikido Security’s reported token spend as a reminder to define the decision before running the test. Teams need to specify which tasks matter, what constitutes a correct answer, how human reviewers will score results, and how the cost of repeated agent activity will be calculated.
They should also test failure behavior. A model that refuses uncertain requests, identifies missing evidence, and avoids taking unauthorized actions may be more useful than one that generates confident but fragile fixes. Evaluations should include repositories and alerts that resemble production conditions, while protecting confidential code and customer data.
For enterprise AI programs, the deployment architecture may matter as much as the model. Teams must compare hosted and self-managed options, logging and retention policies, access controls, tool permissions, and the ability to reproduce a recommendation. These factors can determine whether an AI security system is deployable, even if its raw technical score is strong.
The story also highlights a market challenge. Model performance is increasingly task-specific, while vendors often present results through proprietary tests. Buyers need enough methodological detail to reproduce or challenge those claims. A headline token total draws attention, but it does not by itself identify the best model for a particular security operation.
The most important follow-up is a full account from Aikido Security. Readers will need the list of models tested, the task categories, the evaluation dataset, the scoring criteria, and a breakdown of how the 11.7 billion tokens were allocated.
It will also be important to see whether Aikido Security names a winning model and publishes comparative results, error examples, or cost-per-task data. Independent replication would provide a stronger basis for judging the claim than the current vendor-controlled source record.
Security teams should watch for evidence beyond headline rankings: false-positive rates, patch validation, performance on unfamiliar code, tool-use safety, latency, and the effect of human review. Those measures are more likely to predict whether a cyber AI model can support production work.
Aikido Security’s reported 11.7-billion-token experiment is newsworthy because it frames model selection as an engineering and operating-cost problem, not just a leaderboard exercise. But the available evidence is too limited to support a conclusion about which model is actually best.
The useful lesson for AI builders is methodological: large-scale testing can reveal differences that short demos miss, yet scale alone does not make a benchmark credible. Until the underlying protocol and results are available, the report is best read as a signal of how demanding cyber AI evaluation can become—not as proof of a definitive market winner.
Aikido Security says it spent 11.7 billion tokens comparing cyber AI models, raising questions about benchmark design, cost, and evidence for buyers.