Businesses
Perplexity and Nvidia Are Betting on Local AI as Enterprise Costs and Privacy Come Into Focus
28 Aug 2026

Perplexity’s new Portable Computer moves more AI agent activity onto local Nvidia hardware, pointing to a hybrid future in which routine work runs on-device while powerful cloud models are used only when needed.
For most of the generative AI boom, the model has lived somewhere else.
A user sends a prompt to a cloud data centre, the model processes it remotely and the answer comes back. That setup made sense when the most capable models required infrastructure far beyond what could sit on a desk.
AI agents are beginning to change that equation.
Instead of answering one question and stopping, agents can inspect files, call tools, run code, retry failed steps, generate intermediate outputs and continue working for long periods. That can make cloud inference expensive, particularly when much of the work does not require a frontier model.
Perplexity is now testing a different approach.
On 25 August 2026, the company launched Portable Computer, a local-first version of its Computer agent designed initially for Nvidia DGX Spark and other sufficiently powerful Nvidia GPU systems. The model, agent runtime, conversation state, tools, file processing and security sandbox can all operate on hardware controlled by the user.
Moving More of the AI Agent Onto the Device
Running a local language model is not new. Developers have been doing that for years.
What makes Portable Computer more interesting is that Perplexity is moving much more of the agent architecture itself onto local hardware.
A chatbot mainly needs to accept a prompt and return text. An agent needs an execution loop: it must decide what to do next, gather context, select tools, inspect results and recover when something goes wrong.
Perplexity says its system uses a deterministic orchestrator to control that process, while the language model proposes actions. Approved commands run inside an operating-system-level sandbox rather than receiving unrestricted access to the machine.
That turns local AI from a model experiment into something closer to an operating environment for autonomous software.
Why AI Agents Change the Cost Calculation
The economics are one of the biggest reasons this matters.
A conventional chatbot may generate a few hundred or thousand tokens for a user request. An agent working through a large set of documents can create many more invisible model calls as it extracts data, compares information, checks its own work and produces a final answer.
Under a cloud-only setup, those intermediate steps can all add to inference costs.
Perplexity’s local-first design shifts part of that expense away from metered API usage and towards hardware the company or user already owns. Local inference is not free, of course: hardware still has to be purchased, powered and maintained. But the economics can look very different when a machine is being used continuously.
Nvidia currently lists DGX Spark at $4,699. The system includes 128GB of unified memory and is positioned for demanding local AI workloads, including models far larger than those supported by most consumer GPUs.
For businesses running persistent AI agents throughout the working day, the question increasingly becomes familiar: is it cheaper to keep renting compute, or own more of it?
A Hybrid Model: Local Worker, Cloud Advisor

Perplexity is not arguing that local models can replace the most powerful cloud systems.
Its approach is closer to a division of labour.
The local model handles repetitive or sensitive work, such as reading files, manipulating data and using local tools. If the task requires current web information or significantly stronger reasoning, the system can call an external service or frontier model with user permission.
That creates a simple hierarchy.
The local model becomes the worker. The local orchestrator controls what it is allowed to do. The cloud model becomes an advisor that is called when the problem genuinely requires more capability.
For enterprises, that could reduce the amount of routine work sent through expensive frontier APIs.
Privacy Becomes Part of the Architecture
Local processing also changes the privacy discussion.
Cloud AI providers can offer encryption, retention controls and enterprise security commitments, but organisations still have to decide whether sensitive information should leave their infrastructure in the first place.
If financial records, source code, legal documents or confidential research can be processed locally, they may not need to be uploaded to an external model at all.
Perplexity’s architecture allows local files to remain the primary source while external search or cloud services are used selectively. External capabilities can also be disabled for workflows that need to remain offline.
That could be particularly attractive in finance, legal services, government, cybersecurity and other sectors where data control matters as much as model capability.
Local AI does not remove security problems, however. An agent with access to files, code repositories and enterprise systems can itself become a powerful attack surface if its permissions are poorly managed.
That is why execution controls and sandboxing may prove just as important as the model.
Perplexity Says the Agent Harness Matters as Much as the Model
The company’s own benchmark results also point to another shift in AI development.
On its 53-task Local Knowledge Work Bench, Perplexity reported that Computer running Qwen 3.8 27B on DGX Spark scored 82.6%, compared with 77.6% for Pi and 74.0% for Hermes using the same base model. Perplexity’s post-trained PPLX 27B model scored 85.4%. The figures are vendor-reported and have not necessarily been independently replicated.
The broader point is that useful AI performance increasingly depends on more than model weights.
Tool selection, context management, memory, retrieval, security controls and orchestration all affect whether an agent can complete a task reliably.
For business buyers, that could eventually make the full agent system more important than small differences between standalone model benchmarks.
Why Nvidia Wants AI Running on the Desk
The strategy is also significant for Nvidia.
Local AI does not necessarily compete with its data-centre business. It creates another place where accelerated computing can be sold.
If frontier models remain concentrated in hyperscale data centres, Nvidia supplies the large clusters. If persistent AI workloads move into offices and individual workstations, Nvidia can also supply the hardware closer to the user.
The company describes DGX Spark as suitable for local inference, fine-tuning and long-running autonomous agent workloads, while support for additional GeForce RTX and RTX PRO systems is expected to broaden the market further.
This could eventually create a new category of enterprise equipment: the AI agent appliance.
A legal team might process documents on a local AI server. A finance department could analyse records without sending them outside the corporate network. Developers could use workstation-class agents for coding and testing, while only the hardest reasoning tasks are escalated to cloud models.
The Bigger Shift Is From Sending Data to AI to Bringing AI to the Data
Portable Computer does not mean the cloud is going away.
The largest models will still require enormous data centres, and cloud platforms remain valuable for scale, collaboration and access to the most advanced systems.
But local AI may begin absorbing a larger share of routine execution.
That could change how businesses think about their AI architecture. Rather than choosing between local and cloud systems, organisations may increasingly build a hierarchy in which tasks move between personal devices, internal GPU infrastructure and frontier APIs depending on cost, privacy, latency and difficulty.
The first phase of generative AI largely moved business data towards centralised intelligence.
Local-first agents reverse part of that flow.
They move more intelligence towards the data instead.
For Perplexity, that is a bet on the agent layer. For Nvidia, it creates another market for accelerated computing. For businesses, it raises a more practical question: how much AI work really needs to leave the building?
Source
Share

Sara Srifi
Sara is a Software Engineering and Business student with a passion for astronomy, cultural studies, and human-centered storytelling. She explores the quiet intersections between science, identity, and imagination, reflecting on how space, art, and society shape the way we understand ourselves and the world around us. Her writing draws on curiosity and lived experience to bridge disciplines and spark dialogue across cultures.





