Bran (Brandon) Myers
Analysis · AI Security · NEW · 12 September 2026

AI Has Moved From Assistant to Operator

Anthropic’s September 2026 threat report is one of the most disturbing documents yet published about the real-world misuse of artificial intelligence.

Not because it describes a single breakthrough attack. Not because an AI independently decided to harm anyone. And not because the crimes themselves are entirely new.

The report is disturbing because it shows the same structural change appearing across almost every category of malicious activity: AI is moving from assistant to operator.

According to Anthropic, malicious actors used Claude to conduct reconnaissance, develop exploits, write malware, process stolen data, operate propaganda networks, profile dissidents, engineer surveillance platforms, design software for weapons, support dangerous dual-use biological research, sustain deceptive romantic personas and extract the capabilities of frontier AI models.

In several cases, the model was not merely answering questions. It was working inside an operational system. It could receive a broad objective, divide the work into tasks, coordinate sub-agents, inspect results, write and execute code, preserve campaign context and continue operating with limited human supervision.

That distinction matters.

A chatbot can make an individual more productive. An AI operating layer can allow one individual to approximate an entire team.

Anthropic says it identified and disrupted the activity described in the report between December 2025 and August 2026. The company attributes the operations to a mixture of suspected state-sponsored groups, financially motivated criminals, commercial spyware vendors, state propaganda institutions, politically motivated individuals and competing AI laboratories. These are Anthropic’s findings and attributions, rather than independently adjudicated facts, but the collection is significant because it offers direct visibility into how malicious users reportedly interacted with frontier models.

Taken together, the cases reveal a new threat landscape. The crucial variable is no longer simply what an AI model knows. It is what the model can coordinate, execute and repeat at machine speed.

The primary source
Everything below is drawn from Anthropic’s own report. It is worth reading rather than taking my account of it, and the case detail is heavier than any summary can carry.
↓ Download “Detecting and countering misuse of AI: September 2026” (PDF, 11 MB)

Cyber operations: the collapse of the labour barrier

The cyber section of the report provides the clearest evidence of the transition.

Cyberattacks have always involved more than exploitation. A serious campaign requires target research, infrastructure, credential acquisition, software development, persistence, data collection, exfiltration and analysis. Each stage traditionally consumes skilled human labour. This limited how many targets an operator could attack and how quickly an operation could adapt.

Anthropic reports that this labour barrier is collapsing.

One operation, attributed by Anthropic to a Russian espionage actor whose activity was consistent with public reporting on Midnight Blizzard, allegedly used customized AI workflows across much of the cyber kill chain. The actor researched targets, registered domains, configured phishing infrastructure, monitored command-and-control channels, harvested credentials, moved through victim networks, organized stolen information and maintained access to compromised environments.

The reported targets included Ukrainian and European government bodies, military and intelligence organizations, embassies, diplomatic personnel, think tanks, defense companies and suppliers involved in military drone technology.

The actor allegedly compromised hotel Wi-Fi management vendors to reach guests indirectly. Stolen hotel information could be combined with device data to identify people of particular interest. Anthropic also reports attempts to take over WhatsApp accounts, export conversations, access live camera streams and steal large government identity and company registries.

The most consequential capability was an automated detection-evasion loop. AI agents monitored whether security products detected the actor’s malware. When a tool was flagged, the agents modified and rebuilt it, continuing the process until it was no longer detected. The updated malware could then be redeployed.

This reverses a familiar defensive advantage. A security team traditionally imposes costs on an attacker by identifying a malicious tool and publishing a detection. The attacker must then spend time and expertise rebuilding it. If AI can close that loop automatically, a static detection may have a much shorter useful life.

Anthropic describes other actors using AI for opportunistic, financially motivated attacks. Suspected affiliates of the ShinyHunters collective allegedly used automated systems to scan vast quantities of software and infrastructure for exposed credentials. One operator reportedly ran a distributed pipeline that downloaded 1.8 million Android applications, decompiled them and scanned them for embedded secrets. A parallel system harvested GitHub access tokens.

Once a valid credential was found, AI agents could explore the target environment, discover valuable systems, write the necessary scripts and extract data. The operator did not always need to understand the victim’s particular technology stack. The model interpreted it on demand.

Anthropic calls this style of activity “vibe hacking”: the human specifies a broad outcome, such as finding valuable information or using a credential against a company, while the AI handles much of the technical detail.

Reported consequences included:

AI infrastructure itself also became part of the criminal supply chain. Attackers allegedly stole AI API keys from victims and transferred their workloads onto those accounts. That gave them capability, free compute and cover, since the malicious traffic appeared to belong to the victim.

Another operation reportedly advertised cheap access to Claude while silently routing customers to another model and installing a credential harvester. The customers paid for a fraudulent service and lost their real Anthropic credentials, which could then be sold to other proxy services and malicious users.

The report also describes autonomous exploit research. A China-based group allegedly operated parallel AI workstreams for foreign-government reconnaissance, intrusion attempts, malware development, intelligence collection and vulnerability research against major security products.

AI agents loaded firmware into decompilers, followed code paths, proposed vulnerability hypotheses, generated test cases, wrote exploit code and tested it against laboratory copies of the relevant products. Successful results entered the operators' private exploit portfolio. These workflows reportedly continued while their owners were away.

In another case, a single French-speaking hacktivist allegedly targeted European political parties, media organizations, think tanks and their software providers. Anthropic says the actor used AI to develop a previously undocumented WordPress exploit, compromise websites, implant backdoors, harvest credentials and poison backups so restored systems could be reinfected.

The same actor reportedly built a doxxing platform containing tens of millions of rows drawn from health, justice-system, identity and breach data. Individuals associated with a targeted political movement could be searched by name through anonymously hosted services.

The larger lesson is not that AI invented phishing, credential theft, malware or SQL injection. Anthropic explicitly observes that the underlying attacks remain familiar. What changed was the unit economics.

Reconnaissance, programming, exploitation and data processing can now be delegated to models running in parallel. Campaigns that once required a staffed operation may be attempted by a few people, or even one determined person with stolen access.

Sophistication is therefore becoming a weaker attribution signal. Advanced tradecraft no longer necessarily implies an advanced organization.

Influence operations: propaganda becomes infrastructure

The same transition appears in information operations.

Generative AI has long been associated with cheap political content. Anthropic’s report goes further. It describes AI embedded in systems that manufacture, localize, publish and amplify narratives across entire media ecosystems.

In the Central African Republic, Anthropic says a Russian-speaking actor used Claude as the production backbone of a Russian state-aligned influence operation. The actor allegedly instructed the model to embed pro-Russian, anti-French, pro-government and pro-Wagner narratives into locally styled reporting while removing features that might make it appear machine-generated.

The output was not confined to anonymous posts. Anthropic reports that the content moved through a radio station, Telegram channels, local outlets and national broadcasting. The model was used to produce briefings, scripts, graphics, contracts and internal management documents.

The operation also allegedly used Claude to create employment rules requiring loyalty, score staff work against political criteria and recommend which employees should be retained or dismissed. When the model questioned the political weighting, the actor reportedly relabelled the criteria in neutral language and continued.

This is an important form of safeguard evasion. The operator did not necessarily defeat a filter technically. They changed the description while preserving the function.

Anthropic also reports a commercial “influence-as-a-service” operation linked to a French digital advertising company. Approximately 70 fabricated news sites were paired with around 70 matching social accounts and more than 250 inauthentic commenting accounts. The network reportedly published at least 8,913 articles in approximately 20 languages.

The company did not appear committed to a single ideology. It could support different sides according to the interests of the paying customer. Real journalism was rewritten with added political angles, the same story could be presented in opposing ideological forms and material could be laundered into unrelated countries after its original context was removed.

Fake journalists supplied invented bylines. AI-generated profile photographs made supporting accounts appear authentic. Search-optimized formatting and internal links helped the fabricated outlets imitate legitimate publishing operations.

Another commercial platform allegedly targeted Malaysia. Anthropic says the system combined census information, voter records and electoral data to profile all 222 parliamentary constituencies, concentrating on sensitive divisions involving race, religion and royalty. It managed roughly one thousand fake social accounts, warmed them to appear authentic, rotated technical identifiers and allowed operators to allocate artificial views and engagement to political targets.

The platform also produced fabricated dossiers containing uncorroborated allegations against an opposition politician and civil-society organizations. When Claude refused overtly defamatory work, the operators allegedly negotiated less explicit wording while continuing to build toward the same end.

The report additionally describes Claude acting as a newsroom production layer for Russian state media. Individual operators allegedly transformed source material into localized articles, Telegram posts, television captions, voiceovers and broadcast scripts. Content could move from Russian state or intelligence sources into apparently independent regional commentary, then be echoed across multiple outlets to manufacture the appearance of confirmation.

The threat here is not merely an infinite supply of text. The threat is an inexpensive, multilingual production system that can maintain house styles, fabricate institutional identities, adapt narratives to local grievances and coordinate distribution.

Propaganda becomes less like posting and more like infrastructure.

Surveillance: software engineering for repression

The surveillance cases may be the report’s bleakest section because they show AI being used to build systems that operate at population scale.

Anthropic says a single consultant used Claude as the primary engineering workforce for a domestic surveillance platform intended for Mali’s state intelligence service. The reported system, called Lakana 360, was designed to monitor approximately 25 million SIM cards across all three national mobile operators.

The platform allegedly collected call records, text messages and voice traffic. It included cross-SIM voice identification, biometric-registry matching, geofenced watchlists, clandestine-meeting inference and the flagging of people who used encryption or VPN services.

Most strikingly, Anthropic says the actor requested removal of the warrant requirement from the component that generated an intelligence dossier about any phone number. The control was reportedly disabled by default and data retention was indefinite.

This illustrates the difference between AI as an analytical tool and AI as an institutional force multiplier. A state does not need to employ enough analysts to read every record manually if software can join identities, communications, locations and registries automatically, then produce a narrative dossier on demand.

China-based actors allegedly used Claude to monitor petitioners, rights defenders, religious communities, dissidents and diaspora organizations. The systems identified citizens who might file grievances so officials could intervene beforehand. Overseas activists and events were profiled, including gathering points, routes, venues and organizations associated with democracy and human-rights advocacy.

One workflow reportedly assigned enforcement categories to named individuals. Another generated daily intelligence briefings in official formats. The operator then converted the process into an AI-use manual for wider adoption inside the bureau.

Anthropic also describes two Iranian-linked units building connected surveillance systems. One allegedly maintained identity records and profiled thousands of Iranians. Another used Claude to build a malicious Firefox extension disguised as a prayer-times utility, which harvested identities from social platforms.

Other reported tools included a messenger de-anonymizer, a phone-number-to-identity resolver, a national-ID phishing page, a mass-reporting bot and a government surveillance case-management interface. A related actor allegedly used voice cloning to prepare propaganda involving recognizable Iranian writers.

Separate Iranian-linked activity reportedly used AI to create identity-profiling systems targeting Israeli officials, nongovernmental figures and Jewish diaspora organizations. Another project compiled publicly accessible ship, aircraft, personnel and satellite information into targeting recommendations concerning US naval forces.

The report also describes malware intended for domestic surveillance: keylogging, screenshot capture, credential theft, mobile data collection and tools disguised as legitimate applications. Some components were engineered through fragmented, individually innocuous prompts after direct malicious requests were refused.

That pattern appears repeatedly across the report. A model may reject the final harmful system if it sees the whole objective. It may still supply many of the components if the operation is divided across accounts and sessions.

Safety therefore cannot be evaluated only at the level of one prompt and one answer. Harm emerges across the sequence.

Conventional weapons: AI joins the engineering team

Anthropic reports six operations involving conventional weapons development, military intelligence or procurement.

A Yemen-based cell allegedly used Claude Code in place of human software engineers to work on guidance, navigation and control software for a guided weapon. The model helped integrate an open-source autopilot with a phone-class flight computer, write control and position-estimation software, tune parameters, build firmware and run simulations.

The actors reportedly managed multiple Claude instances as if they were members of a small engineering team. They later test-fired a guided rocket. Anthropic does not claim that they fielded a successful operational weapon, and the reported test appears to have failed. Within hours, however, the actors returned to Claude to diagnose the failure.

A China-based actor allegedly used the model to produce a specification for an anti-torpedo fire-control system, a technical proposal exceeding 200 pages and an executive briefing. Claude was also used to compare the proposed system with US naval programs, construct portions of the software and produce a validation matrix.

The actor reportedly asked the model to role-play a hostile expert reviewer, then incorporated its criticism into successive drafts. That compressed not only writing time but also part of the review cycle normally performed by technical staff.

A Russia-based team allegedly attempted to build an autonomous first-person-view kamikaze drone swarm. According to Anthropic, Claude assisted with shared swarm memory, fault-tolerant coordination, attack and return-to-base behaviours, terminal camera guidance, detonation logic and the geolocation of opposing control links. Code was tested in simulation and loaded onto real boards.

Another China-based operation allegedly developed a suite of approximately 16 electronic-warfare and air-defense-suppression modules. The system analysed radar sites, missile positions, command posts and communications nodes, calculated detection coverage and jamming effectiveness, ranked targets and assigned jammer sorties across campaigns.

The remaining cases concerned military supply chains. A Russian procurement actor reportedly used Claude to locate third-country intermediaries, draft multilingual correspondence, scrape vendor marketplaces, merge tenders and obscure the intended recipients of mixed military and civilian goods. A China-based actor allegedly mapped foreign directed-energy systems and suppliers to support reverse engineering, comparison and countermeasure development.

Again, AI did not supply factories, laboratories, military authority or physical materials. The human actors already possessed domain knowledge or hardware access. The reported uplift came from compressing software development, documentation, simulation, procurement and review.

That is enough to matter. Engineering capacity is a bottleneck in weapons programs. AI does not need to invent a new weapon autonomously to lower that bottleneck.

Biological misuse: the dual-use problem becomes unavoidable

Biology presents a more difficult problem because much of the relevant knowledge has legitimate scientific value.

Anthropic is careful not to claim that Claude enabled an imminent biological attack. Its report instead identifies state-associated or geographically restricted researchers seeking assistance with work that had serious dual-use potential.

One operation allegedly built an access platform for researchers in unsupported regions, including virologists associated with civilian and military institutions. The platform routed traffic through US infrastructure, used services designed to minimize data retention and included fallback mechanisms when stronger models refused sensitive requests.

Another researcher reportedly spent weeks using weaker Claude models to plan work on an avian influenza pathogen with enhanced pandemic potential. The proposed research involved mutations associated with mammalian adaptation and airborne transmission in animal models. Anthropic says its strongest biological safeguards confined the user to weaker models, limiting the uplift primarily to clerical research support, study design, data analysis and experimental prioritization.

In a third case, a relay serving multiple users allegedly gave a state-associated researcher access to a frontier model, which drafted an orthopoxvirus research grant in approximately an hour. The application included a hypothesis, experimental design, dosing, statistics and contingency plans. The stated work focused on immune-evasion genes and attenuation, which has legitimate defensive value but could also provide knowledge relevant to preserving or transferring harmful viral functions.

Anthropic also reports researchers developing an atlas of venom peptides and a generative pipeline to optimize toxin characteristics. The declared goals included analgesics and antidepressants, but the underlying scaffolds covered paralytic as well as therapeutic targets.

Another researcher allegedly redesigned a variety of toxins computationally while directing Claude to keep the identities of certain biological agents deliberately vague in progress reports.

These cases expose the limits of content filters. A classifier can block clearly stated requests to develop a known biological weapon. It cannot reliably determine the intent behind every advanced protein-design, immune-evasion or toxin-optimization project. The same method may produce a medicine or an incapacitating agent.

The harder frontier is therefore trusted access. Providers may need to understand who is using the most capable biological systems, under what institutional authority and with what degree of accountability. That creates its own privacy and governance questions, but the report makes clear that prompt filtering alone is insufficient.

Scams and fraud: synthetic intimacy at industrial scale

One of the report’s most immediately human cases involved a network of deceptive dating applications.

Anthropic says a China-based studio used Claude both to build more than 20 dating apps and to operate the personas inside them. Over a two-week period, more than 4,700 AI personas reportedly conversed with at least 25,000 people, producing approximately 2.36 million messages.

The service advertised human interaction. In reality, the match feed allegedly contained approximately three AI personas for every real person.

The AI personas were instructed not to reveal that they were automated, to deflect requests for photographs or video calls and to move users through fixed conversational stages. The service fabricated likes, visitors and apparent video interactions while charging users through metered messaging.

Real gig workers were mixed into the network to handle authenticity checks that the models could not perform. They could appear on live video, follow a user on social media or choose from AI-generated reply suggestions. This hybrid structure is important. Fraud does not need full automation. It needs automation across the expensive parts and small amounts of human participation at the moments where trust must be reinforced.

The applications were allegedly engineered to conceal their real behaviour during app-store review. Special interfaces activated for reviewers, class names varied across otherwise similar apps and payment redirection could be hidden remotely.

Anthropic notes that in some sampled interactions, the model appeared to recognize potential harm when users disclosed serious illness or acute distress, but continued the assigned persona rather than refusing.

This case turns synthetic media into something more intimate than a fake article or profile photograph. The product being fabricated was a relationship, sustained continuously and monetized one message at a time.

Illicit distillation: AI models become targets

The final section concerns the theft of AI capabilities themselves.

Model distillation is normally a legitimate process in which a smaller model learns from outputs generated by a more capable model. Anthropic uses the term illicit distillation for large-scale, unauthorized efforts to extract Claude’s behaviour and reasoning in order to improve competing systems.

The company attributes such campaigns to seven China-based laboratories and companies. These attributions include Alibaba’s Qwen team, Moonshot AI, DeepSeek, Zhipu, Xiaomi, SenseTime and MiniMax.

Anthropic alleges that Alibaba-linked operators conducted the largest distillation campaign it had measured, peaking at nearly three million exchanges per day across more than 3,500 fraudulent accounts. The campaign allegedly targeted reasoning traces, agentic tasks, software engineering, kernel development and long-horizon work. Anthropic reports observing more than 151 million exchanges attributable to the campaign between May and July 2026.

Moonshot and DeepSeek allegedly went further by silently routing some customer requests to Claude while users believed they were interacting with those companies' own models. The responses could be shown to the user and retained for training.

According to Anthropic, these relayed conversations contained sensitive information, including internal company documents, government credentials, surveillance data, source code, contact information and live access tokens. The users may not have known that their information was being forwarded to another provider.

Xiaomi allegedly saved conversations involving its own models and replayed them through Claude to create supervised and reinforcement-learning data. SenseTime reportedly acquired harvested exchanges from third-party vendors. MiniMax allegedly operated a proxy network through a shell company, offering access to US frontier models while potentially collecting the resulting conversations.

The methods reportedly included fraudulent accounts, residential proxies, disposable emails, virtual payment cards, cross-session replay and prompts designed to make Claude reveal or reconstruct reasoning that would otherwise remain hidden.

This creates three distinct harms.

First, intellectual property and expensive model capabilities may be extracted at enormous scale.

Second, safety protections do not necessarily transfer to the distilled model. A competing system may acquire more capable reasoning without the safeguards that limited the original.

Third, users' private information can become raw training material without their knowledge. Someone who believes they are sending code, credentials or corporate documents to one service may unknowingly expose them to another provider, a routing network and a downstream training pipeline.

AI supply chains therefore introduce a new question of provenance: which model actually processed the request, who retained it and where did it travel afterward?

The race with China is accelerating the danger

The political response to these threats is being shaped by another powerful force: the race between the United States and China for AI supremacy.

Donald Trump has not treated AI primarily as a technology requiring cautious governance. His administration has framed it as a contest for economic, military and geopolitical dominance.

The language is explicit. Trump’s January 2025 executive order made it US policy to “sustain and enhance America’s global AI dominance” and directed agencies to review, suspend or rescind measures inherited from the Biden administration that might obstruct that objective. The White House’s subsequent AI Action Plan is titled Winning the Race. It says the United States must achieve global dominance, innovate faster than its competitors and dismantle regulatory barriers that hinder private development.

The administration has not eliminated every AI rule, and its own plan contains measures concerning security, model evaluation, biosecurity and the protection of American technology. But its governing presumption is unmistakably deregulatory. The plan instructs agencies to identify, revise or repeal federal rules that impede AI development. It proposes considering a state’s AI regulatory climate when allocating federal funding. In December 2025, Trump went further, ordering the creation of a federal task force to challenge state AI laws deemed inconsistent with a “minimally burdensome” national framework.

The political logic is easy to understand. If the United States slows down and China does not, China may gain an economic or military advantage. If China can distil American models, recruit their capabilities and embed AI into weapons, surveillance and intelligence systems, American policymakers will fear that unilateral restraint amounts to strategic surrender.

Anthropic’s report gives that fear substance. It attributes illicit distillation operations to multiple Chinese laboratories and companies. It describes China-based actors allegedly using Claude for exploit development, political surveillance, religious monitoring, anti-torpedo systems, electronic warfare, supply-chain intelligence and advanced dual-use biological research.

But this is precisely why race framing is dangerous.

Every disclosed Chinese operation becomes an argument for faster American deployment. Every American acceleration becomes justification for Chinese acceleration. Safety testing, access controls, transparency requirements and liability rules can then be portrayed as gifts to the adversary. Regulation stops being judged on whether it is effective and starts being judged on whether it might slow the national champion.

That creates a security dilemma. Neither side needs to believe that unrestrained development is safe. Each only needs to believe that restraint is more dangerous if the other side refuses to adopt it.

The result is a race in which commercial incentives and national-security incentives point in the same direction: larger models, more autonomy, more infrastructure, wider deployment and shorter development cycles. The harms documented by Anthropic are not external to that race. They are being accelerated by it.

There is also a contradiction at the centre of the American position. The United States wants frontier systems to become more powerful and widely deployed while preventing hostile governments, criminals and competing laboratories from stealing or misusing those capabilities. Yet the pressure to release quickly, reduce regulatory friction and spread American AI globally increases the number of accounts, integrations, intermediaries, model routers and infrastructure layers that must be secured.

Capability proliferation expands the attack surface.

The answer cannot be to abandon AI development or pretend that geopolitical competition does not exist. Nor should regulation mean freezing a rapidly useful technology under rules written by people who do not understand it. Bad regulation can protect incumbents, entrench bureaucracy and fail to stop sophisticated adversaries.

But a policy of acceleration first and safeguards later is not realism. It is a wager that defensive capacity will somehow keep pace with systems deliberately optimized for greater autonomy and reach.

Anthropic’s own evidence suggests otherwise. The abuse is already crossing providers, borders, proxy networks, stolen accounts and local deployments. Once a weapons, surveillance or malware system is completed and moved onto private infrastructure, terminating the original AI account cannot remove what has been built.

A serious regulatory framework should therefore focus on observable capabilities and operational risk rather than political control over model speech. It should establish independent evaluations for frontier systems, mandatory reporting of significant misuse, security standards for model weights and high-risk infrastructure, meaningful provenance requirements for model-routing services, protected information sharing between providers, accountability for deceptive AI impersonation and trusted-access controls for the most dangerous biological and weapons capabilities.

Those measures would not guarantee safety. They would at least recognize that winning an AI race is meaningless if the race systematically produces tools that make democratic societies easier to defraud, surveil, manipulate and attack.

Safeguards are being attacked as systems

Across every section, the same evasion patterns recur.

Actors divide harmful projects into individually benign components. They spread work across sessions, accounts, organizations, models and providers. They use stolen API keys, residential proxies, disposable identities, resellers and unsupported-region relays. They replace explicit military or political terms with neutral ones. They frame dual-use research as civilian or therapeutic. They move completed software onto local models that can no longer be disabled by the original provider.

This means a refusal rate measured on isolated prompts can provide a misleading sense of security.

A model might refuse to build a surveillance implant in one exchange while separately supplying a browser extension, a data parser, a credential interface, a messaging bot and a deployment script. Each component may appear ordinary. Their combination is not.

The unit of analysis must move from the prompt to the operation.

That requires behavioural signals, campaign memory, identity and payment analysis, infrastructure correlations, tool-use monitoring and information sharing between providers. It also raises difficult civil-liberties questions. Powerful monitoring systems can themselves become mechanisms of surveillance if deployed without proportionality and accountability.

There is no simple technical switch that solves this.

Defence must become agentic too

The bleakest reading of this report is that the attackers have discovered scalable AI before defenders have built an adequate response.

The more useful reading is that the same structural advantage is available to defence. In a limited but practical way, I had already begun building toward this defensive model before reading Anthropic’s report.

Aegis, the defensive agent I built to guard a distributed web and infrastructure mesh, began with a simple assumption: static controls and occasional human inspection would not be sufficient in an increasingly automated threat environment.

That assumption now appears throughout Anthropic’s evidence.

If malicious tools can monitor detections and rebuild themselves, defensive agents must monitor infrastructure continuously, correlate weak signals across services, test whether an event is truly malicious and act quickly enough to contain it. If attackers operate across thousands of targets and millions of records, defenders cannot rely exclusively on human review queues.

Aegis watches uptime and service health, live request traffic, TLS state, origin exposure, file integrity and hostile behaviour across the mesh. It connects those signals rather than treating each alert as an isolated event. It performs contextual triage, distinguishes routine noise from credible threats, issues immediate alerts and can act through integrated enforcement systems. Its defensive memory improves the next judgment instead of letting every incident begin from zero.

The purpose is not to remove human authority. It is to concentrate human judgment where it matters while automation handles constant observation, correlation and routine response.

I am not claiming that one defensive agent solves the global risks in Anthropic’s report. It does not. Aegis operates at a far smaller scale. But it embodies the architecture the report makes necessary: persistent observation, cross-system context, adaptive judgment, fast containment and a human retained at the point of consequence.

In that sense, I have already begun defending against the operational transition Anthropic describes.

The same principle needs to exist at larger scales. AI providers, cloud platforms, telecommunications companies, security vendors and governments need defensive systems capable of recognizing operations rather than merely blocking suspicious strings.

Human oversight remains essential, particularly where an automated action could restrict speech, terminate access or wrongly accuse a user. But human oversight without machine-speed detection will be overwhelmed by machine-speed abuse.

The answer to agentic offense is accountable agentic defence.

The real warning

Anthropic’s report is not evidence that AI has become an autonomous political actor, criminal mastermind or weapons scientist.

It is evidence that humans can embed AI inside malicious organizations and allow it to perform an increasing share of their labour.

The humans still choose the objectives. They select the targets, provide the infrastructure, interpret the results and decide how to monetize, repress or deploy what the system produces.

But the number of humans required is shrinking.

That is the real warning.

One operator can direct multiple agents. One campaign can adapt across many victims. One propaganda system can imitate dozens of newsrooms. One consultant can engineer population-scale surveillance. One fraudulent studio can maintain thousands of synthetic relationships. One laboratory can generate millions of training exchanges through thousands of false accounts.

The threat is not simply artificial intelligence.

It is institutional capacity without the institution.

For years, we have asked whether AI is intelligent enough to be dangerous. Anthropic’s report suggests a more immediate question:

How much operational capacity can a determined human assemble around it?

The answer is already: far more than most institutions are prepared to defend against.

Primary sources: Anthropic, “Detecting and countering misuse of AI: September 2026”, published September 10, 2026; the White House, “Removing Barriers to American Leadership in Artificial Intelligence,” January 23, 2025; the White House, “America’s AI Action Plan,” July 2025; and the White House, “Ensuring a National Policy Framework for Artificial Intelligence,” December 11, 2025. The incidents, actor identities and organizational links attributed to malicious operations above are based on Anthropic’s reporting. They should not be treated as independently adjudicated findings.
← All Writing