Blog

AI is Escaping the Cloud

Loading the Elevenlabs Text to Speech AudioNative Player...

Cloud AI isn’t disappearing—it’s spreading. Discover how edge AI, local agents, robots, and autonomous systems will reshape business, expand the AI market, and lead to the rise of the digital employee in a box.

AI is Escaping the Cloud

Why AI is moving to the edge and turning machines into ‘digital employees in a box’

For the first few years of the generative AI revolution, intelligence seemed to live somewhere far away.

We typed a prompt into ChatGPT, Claude, or Gemini. It disappeared into the cloud, thousands of GPUs sprang into action, and an answer came flying back. AI felt less like software and more like summoning an oracle.

That made sense (then). The models were enormous and only a handful of companies could afford to run them. But AI’s on the move, spreading into corporate facilities, factories, vehicles, robots, workstations, laptops, phones, cameras, and appliances—not because the cloud has failed, but because intelligence increasingly works better near the data, user, or machine taking action. Leaders need to understand that the architecture of enterprise infrastructure is shifting.

The next era won’t be cloud or edge. It’ll be cloud and edge: a distributed fabric of intelligence extending from enormous AI factories to the devices in our pockets—and perhaps eventually to computers orbiting above us, a further expansion of the cloud into the heavens.

AI Is Shrinking Faster Than Its Capabilities Grow

Large models ignited the current AI boom, producing breakthrough capabilities in language, coding, reasoning, and multimodal understanding. But bigger isn’t always better. As I’ve discussed previously, most business tasks need a model that’s smart enough, fast enough, reliable enough, and affordable enough for the job. But not more.

Distillation, quantization, and pruning

Several forces are at work shrinking AI models. Distillation allows a smaller model to learn from a larger one. Quantization reduces the numerical precision used to represent a model, dramatically lowering its memory and computing requirements. Pruning removes weights and connections that contribute little to performance. Better architectures and algorithms allow newer models to accomplish more with less. Put this all together and models are shrinking rapidly for the same level of intelligence.

Researchers use a variety of techniques to shrink models while retaining capability

Edge hardware is gaining AI acceleration

At the same time, AI accelerators are appearing in PCs and phones, memory is improving, and inference software is becoming more efficient. Model routers can send a simple task to a small local model and escalate a difficult one to the cloud.

The amount of computing required to deliver a useful unit of intelligence is falling, making it possible to add intelligence to smaller and smaller devices.

Mainframes didn’t disappear when PCs arrived, and PCs didn’t vanish when the smartphone came. Computing constantly expands into new places and markets. AI is doing the same.

Why Intelligence Wants to Live Near the Work

Moving AI to the edge isn’t simply a technical achievement. In many cases, it’s a business necessity.

Privacy and control

The most valuable business data is often the data companies are least willing to send elsewhere: customer records, legal documents, product plans, source code, financial information, medical data, or proprietary operating knowledge.

Running models locally or on-prem keeps that information inside the firewall and gives companies more control over model choice, updates, retention, and auditability.

Local doesn’t automatically mean secure; a poorly governed internal agent can still cause damage. But local infrastructure can create stronger boundaries and more predictable data handling.

Speed and reliability

Cloud AI is fast, but physics still takes its toll. Data has to travel to a data center, be processed, and return. For writing an email, a slight delay hardly matters. For a robot trying not to collide with a person, or a vehicle deciding whether the shape ahead is a plastic bag or a child, it matters enormously.

These systems must be resilient and keep operating even when a connection becomes weak or unavailable. A self-driving vehicle can’t pull over every time it loses 5G, and a humanoid robot can’t freeze while shaking my well-earned martini just because the cloud’s having a bad day.

For autonomous machines and robots, the cloud is great for training models, distributing updates, coordinating fleets, and analyzing what happened after the fact. But machines must have enough intelligence aboard to perceive, decide, and act safely.

Economics

Cloud services make AI a variable cost, which is wonderful when demand’s uncertain. But once a workload becomes constant, renting intelligence by the token feels like taking a taxi to work every day. At some point, owning a car is probably much cheaper.

As organizations deploy thousands of agents, making millions of routine decisions, AI may become a new class of capital equipment. Companies won’t only subscribe to intelligence; they will own machines that produce it.

The Edge Is Not a Place. It's a Continuum.

When people hear “edge AI,” they often imagine a model running on a smartphone. That’s part of it, but the edge is a much broader landscape.

A large enterprise might operate private AI infrastructure inside its data center. A department could share an AI workstation running multiple digital employees, a small business could host agents in a server closet, and have several robots working in the warehouse, each with onboard intelligence.

The most capable enterprise systems will move work dynamically across this continuum. For example, an agent running on equipment inside the firewall might use a local model to handle private documents or routine tasks, then call a larger cloud model when it encounters something unusually complex.

Physical AI Will Pull Intelligence to the Edge

Nothing makes the edge case clearer than AI entering the physical world.

Amazon has already deployed more than one million robots across over 300 facilities. Its DeepFleet foundation model coordinates their movement like an intelligent traffic-control system and is expected to improve fleet travel efficiency by 10 percent. Amazon’s next-generation fulfillment center in Shreveport uses ten times more robotics than its traditional facilities.

Amazon has serious automation ambitions

Amazon hasn’t published a reliable forecast for its future robot count, but the trajectory is clear. Reports based on internal documents suggested that the company hoped to automate as much as 75 percent of its operations by 2033 and avoid more than 600,000 additional hires, although Amazon disputes that perspective. The exact number is less important than the direction of travel: the world’s largest logistics operation is becoming a vast, distributed machine-intelligence system.

Billions of humanoid robots by 2040?

Humanoid robots will expand intelligent automation far beyond warehouses. Goldman Sachs estimates annual humanoid shipments could grow from roughly 20,000 in 2025 to more than 250,000 in 2030 and 1.4 million in 2035. Morgan Stanley’s longer-range scenario suggests nearly one billion humanoids could be in use by 2050. Other forecasts point to far higher numbers still, perhaps as many as 10 billion robots by 2040. Bank of America says 3 billion humanoids by 2060.

Bank of America Global Research forecast for humanoid robot owned units, split by sector

Even a fraction of that growth would create enormous demand for edge chips, memory, sensors, models, networking, simulation, and security.

Robotaxis push even more AI to the edge

Autonomous vehicles will add another wave. Goldman Sachs forecasts that the worldwide commercial robotaxi fleet could grow from roughly 7,000 vehicles in 2025 to around one million in 2030 and six million in 2035. Every vehicle will be a rolling supercomputer, perceiving its surroundings and making a continuous stream of safety-critical decisions.

The more AI takes action in the real world, the less practical it becomes to place intelligence somewhere else.

The Rise of the Digital Employee in a Box

The edge story is also coming to the office.

Apple’s Mac Mini and Mac Studio have become popular machines for running local models and always-on agents. Their unified memory architecture allows the processor and graphics cores to share a large pool of memory, making it possible to run AI models that might not fit into the dedicated memory of a conventional consumer graphics card. Some of the latest AMD-based PCs have a similar capability and Intel is expected to follow suit.

OpenClaw agents saw Mac Minis fly off the shelf

Open-source frameworks such as OpenClaw and Hermes Agent have encouraged enthusiasts to turn PCs into assistants, coding agents, and automation servers. During 2026, higher-memory Mac Minis developed long shipping delays. Business Insider linked some demand to the OpenClaw craze, while later reporting pointed that a memory shortage may also be the culprit. Apple hasn’t confirmed that agents caused the shortage, but have hinted as much.

The important takeaway: The PC is evolving from a machine on which a person performs work into a machine on which digital workers perform work.

Today, someone might dedicate a Mac Mini to a personal assistant. Tomorrow, a small business might buy an appliance containing a researcher, coder, bookkeeper, and customer-service agent. A department could operate hundreds of agents on shared local infrastructure.

Nvidia is seizing the moment

NVIDIA is designing directly for this possibility. In partnership with Microsoft, it introduced RTX Spark, a platform for Windows PCs built around personal agents. NVIDIA says RTX Spark systems will offer up to 128GB of unified memory, a petaflop of AI performance, and enough capacity to run 120-billion-parameter models locally. If you’re not into the tech stats, that’s ok. Your takeaway should be that new PC hardware is impressively specified to run local AI models and agents.

At the other end of the desk, NVIDIA’s forthcoming DGX Station for Windows is designed to run models of up to one trillion parameters and hundreds of agents simultaneously. This isn’t really a PC in the traditional sense. It’s departmental AI infrastructure wearing office clothes.

What we’re seeing here is the creation of a fascinating new category: the digital employee in a box.

NVIDIA Just Validated the Bigger Thesis

NVIDIA recently reorganized its business reporting around two market platforms: Data Center and Edge Computing.

Its new edge category includes PCs, workstations, telecom infrastructure, robots, and automotive systems supporting agentic and physical AI. In its first quarter of fiscal 2027, NVIDIA reported $6.4 billion in Edge Computing revenue, up 29 percent from a year earlier. A company like NVIDIA doesn’t redraw its financial map for a market it considers peripheral. They clearly expect edge to become a significant part of future business. They still reported an astonishing $75.2 billion in quarterly Data Center revenue, so this isn’t evidence that cloud AI is fading. It’s an indication that the next layer of growth will come from intelligence spreading beyond the cloud and into many types of machines.

The Cloud Is Expanding Too—Perhaps All the Way into Space

The cloud remains essential. Frontier models require enormous training clusters, businesses need elastic capacity, and fleet coordination, simulation, and global learning systems will continue to depend on hyperscaler infrastructure. Many countries are building sovereign AI capacity to guard against geopolitical risk (the recent Claude Fable 5 shutoff spooked them all), so data centers are spreading geographically, even as intelligence becomes more distributed locally. So, AI isn’t shifting from one bucket to another. The bucket is getting bigger.

Next Stop—Orbit

And the next step is orbit. NIMBYism, collapsing launch costs, limited terrestrial energy, and simple economics are pushing AI data centers into space. There’s a delightful paradox here. An orbital data center serving Earth might be considered part of the cloud. But for a satellite producing enormous quantities of imagery, it is the edge.

The distinction is becoming less about altitude or ownership and more about proximity to the data and the action.

Business Leaders Need an Intelligence-Placement Strategy

Most companies have spent the past few years asking which model they should use. The next question is where each kind of intelligence should live.

This decision should reflect six key factors: required capability, data sensitivity, desired response time, offline resilience, economics, and the control needed over updates, behavior, and governance.

Some workloads belong in the cloud. Some belong on prem. Some belong in a workstation, vehicle, robot, or device. Many tasks will move across several locations during a single task.

Business leaders should resist locking their AI strategy to one model, provider, or computing environment. The winning architecture will be hybrid, flexible, and designed to route intelligence to the right place at the right moment.

Intelligence Is Becoming Infrastructure

The first phase of generative AI taught us to summon intelligence from the cloud. The next phase will embed it throughout the world around us.

Some intelligence will live in enormous AI factories. Some will sit under the desk inside offices and homes. Some will be embedded in vehicles, robots, and everyday products. Some may eventually orbit above us.

Cloud AI will continue to expand rapidly—it’s currently doubling in capacity every 7 months. And it will soon be joined by an enormous new layer of distributed intelligence, an expansion that will increase demand for AI silicon, memory, networking, models, infrastructure, security, and applications—and make intelligence more private, responsive, resilient, and more deeply integrated into every business.

The AI buildout has only just begun.

Artificial Wisdom

The unlimited curated collection of resources to help you  get the most out of AI

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.

#1 AI Futurist
Keynote Speaker.

Understand what AI really means for your business and how to build AI-first organizations. Get expert guidance directly from Steve Brown.

Former Exec at Google Deepmind & Intel
Entrepreneur and Acclaimed Author
Visionary AI Futurist
AI & Machine Learning Expert