Always-on agents just hit the limits of compute, cost, and control

by Tony Thomas
The OpenClaw situation is a warning shot for enthusiasts of always-on agents.
On April 4, 2026, Anthropic changed how Claude subscriptions work with third-party harnesses such as OpenClaw. They no longer cover unlimited use of agents on flat-rate plans. Sure, you can still run them, but now you must pay through API billing or usage bundles. Anthropic said the increased use of OpenClaw and similar harnesses placed strain on their systems and did not resemble normal subscription traffic.
Why OpenClaw Became So Popular
OpenClaw and its derivatives took off because they turn models into agentic workers. They run locally, connect to frontier models, and execute tasks across apps and time zones. The emphasis is on persistence rather than chat or one-off prompts. Even when you are sleeping, your squad of agents keeps working in the background. That autonomy makes it expensive.
While a chat stops, an agent keeps going. It plans, retries, summarizes, calls tools, and loops. Left alone, it keeps burning tokens at a steady rate. Flat subscriptions were built for short bursts, not waves of activity, so the mismatch is inevitable.
The OpenAI Hire Changes the Context
In February 2026, OpenClaw’s creator, Peter Steinberger, joined OpenAI to work on personal agents. Around the same time, OpenClaw moved to an independent open source foundation supported by OpenAI. Weeks later, Anthropic removed OpenClaw and similar harnesses from subscription coverage. The timing adds context, and the sequence is hard to ignore.
Third-party harnesses sit comfortably between users and models. They decide when to call the model and how long to loop. That makes demand unpredictable. Metered pricing restores control. It also shifts profit back to the labs.
Subscription Pricing vs Agent Reality
Of course, this extends beyond OpenClaw. Other third-party harnesses still work, but they no longer ride on flat plans, and usage is metered. Continuous agent use now requires continuous payments.
Agent workflows scale with time. Leave them running, and they keep consuming compute. Tools that manage email, calendars, and tasks generate a steady load. Infrastructure built for interactive requests handles spikes. It struggles with constant demand. And there is only so much compute to go around.
Flat pricing assumes pauses. Agents remove pauses. That is the core conflict.
The Reaction From the Community
Part of what pushed me to write this was the reaction from developers building agent-driven workflows. A video from Network Chuck captured the frustration. He described receiving notice that subscription coverage for tools like OpenClaw was ending. Many users chose Claude for that flexibility.
Chuck also suggested the move came down to infrastructure strain or subsidized usage. His tone was jarring. These are no longer experiments. People are wiring agents into daily systems. Pricing changes are a gut punch. And it puts strain on people’s businesses and their lives.
Platform Control Is the Bigger Story
This also looks like a shift toward increased platform restraint. Labs are not just shipping models anymore. They are building full agent stacks that include orchestration, tools, scheduling, and memory. The trend is now in creating packaged solutions rather than just doling out compute.
Third-party harnesses sit outside that stack. They drive loops that the provider can no longer control. Moving orchestration inside the platform limits runaway usage and stabilizes demand. It also centralizes the agent layer.
Security and Autonomy Tradeoffs
Autonomous agents often need wide system access. They read files, trigger actions, and move across apps. That expands risk. Prompt injection becomes more dangerous. Automation mistakes scale faster.
Data exposure increases. And runaway agents can burn tokens and suck up compute at alarming rates. Provider-controlled agents reduce some of that surface area. Less freedom. More guardrails.
The Local Hardware Escape Hatch Is Closing
Local inference once looked like the perfect workaround. You could buy hardware, run open models, and skip token costs.
That window narrowed quickly. GPU demand surged. High-VRAM cards became expensive. Plus, larger setups require power, cooling, and upkeep. Lots of power. Even modest rigs add ongoing cost.
As a result, local inference stops looking like a lower-cost alternative. It becomes a large capital expenditure plus ongoing operating expenses. Continuous agents amplify both.
Between Cloud Tokens and Local Hardware
This leaves an awkward middle path.
Cloud models get expensive when they run constantly, and local hardware that can replace them is also expensive. Neither side offers a cheap solution.
Agentic workflows push requirements higher. Longer context. More frequent calls. Increased tool usage. Small local models can route or filter, but full autonomy still requires frontier models.
So the compromise is a hybrid. You run lightweight steps locally, call frontier models when needed, add budgets, and limit loops. It works, but autonomy shrinks.
Using fully local agents becomes harder. Cheap cloud agents are fading. That is the dilemma.
The End of Cheap Unlimited Agents
This is not the end of agentic AI. It is the end of cheap, unlimited agents.
OpenClaw still runs. It just pays for what it uses. If it runs all day. And the meter runs all day.
That pushes designs toward boundaries such as budgets, step limits, and event-driven triggers.
Hybrid stacks become more common, with smaller models handling planning and frontier models handling heavier reasoning.
The goal shifts from autonomy to efficiency.
What Happens Next
The OpenAI hire underscores how central orchestration has become. Agentic control is now the center of the red-hot competition between labs.
OpenClaw did not fail. It just exposed the economics. Always-on agents turn models into infrastructure. Infrastructure costs money. Subscriptions blur that cost. Metering makes it visible.
Expect more tightening, higher prime time rates, fewer open loops, more bounded agents, and more hybrid setups.
Letting agents run forever in the background at a massive scale was always untenable. And the math finally caught up.

