
Perplexity is launching Portable Computer as we speak, a model of its agentic “Computer” platform that runs solely on {hardware} customers already personal — beginning with Nvidia’s DGX Spark desktop supercomputer and Linux machines geared up with Nvidia RTX GPUs.
The launch, developed in shut partnership with Nvidia, is considered one of the most aggressive makes an attempt but to transfer critical AI agent workloads off the cloud and onto native units. The mannequin, the person’s information, and the work itself can all keep on the machine. Work accomplished domestically consumes no billing credit, and the firm says each job begins on the machine by default — with the system asking permission before sending any particular person step to a extra highly effective frontier mannequin in the cloud.
“We have principally introduced the very same UI to a completely native app,” stated Nate, Perplexity’s vp of engineering for infrastructure and enterprise, throughout a press briefing Monday. “This incorporates the entirety of the agent harness and inference and all the things wanted to do work domestically.”
For Nvidia, which has spent the previous two years promoting the world on trillion-dollar AI information facilities, the announcement alerts one thing subtler however strategically necessary: the chipmaker believes native AI has crossed a threshold from hobbyist curiosity to sensible software — and it desires to promote the {hardware} that runs it.
“Native AI reached an inflection level,” stated Nader, Nvidia’s director of developer expertise, who focuses on developer tooling and open supply. “For the longest time, it was hobbyists and lovers, and so they have been operating these quantized fashions that have been quantized down to be tremendous tiny… And whereas that is cool, it is not tremendous sensible. However all that modified with lots of these new open supply fashions which have come out that are tremendous helpful.”
How Transportable Laptop packages a full native AI stack right into a single app
Perplexity Computer, the firm’s agentic platform for data work, orchestrates AI fashions, information, instruments, and internet entry to full multi-step duties — reviewing folders of paperwork, analyzing information, producing reviews, and pushing outcomes into enterprise programs. Portable Computer replicates that have domestically: the native fashions, agent harness, inference engine, instruments, app connectors, and a safety sandbox come packaged collectively in a single system. That bundling is the level. With most native AI stacks as we speak, customers should assemble and function these items individually — downloading mannequin weights, standing up an inference server, wiring collectively instruments, and tuning efficiency.
“Traditionally it is simply been actually painful to carry up the native AI stack,” Nate stated. “With Transportable Laptop, we actually targeted on simply making this a very simple expertise the place you possibly can stand up and operating in a short time.”
In a single demo Monday, the system performed the function of a retail investor reviewing a folder of 1099s and funding paperwork — the form of delicate monetary materials many customers would hesitate to add to a cloud service. Working a 27-billion-parameter Qwen model at full GPU utilization on a DGX Spark, the agent reviewed every doc and flagged circumstances the place the hypothetical investor was paying pointless charges. The interface ingredient that usually shows a operating tally of cloud credit “is simply parked at zero,” Nate famous, “as a result of all of this is occurring on the machine.”
A second demo confirmed the hybrid aspect of the product. Enjoying a startup founder, Nate requested the agent to analyze a CSV of person funnel information domestically, then push the completed evaluation to a Slack channel utilizing Perplexity’s connector ecosystem — proof that local-first does not imply disconnected.
The system additionally connects to Google Drive, Gmail, and GitHub, and may escalate to a frontier cloud mannequin when the native mannequin hits its limits. At launch, customers can arrange Qwen 3.8 27B or PPLX 27B, a model Perplexity has post-trained on its personal harness, with Nvidia’s Nemotron 3.5 Lightning coming quickly.
Transportable Laptop arrives as we speak for Pro, Max, Enterprise Pro, and Enterprise Max subscribers on Linux, with Home windows help following in September. Any RTX GPU with at the very least 24GB of VRAM — roughly a GeForce RTX 3090 or newer — clears the bar, a threshold Nate known as “kind of the flooring the place we actually need to be sure that we will ship a terrific expertise, however steadiness that with making it broadly out there.”
Why co-designing the mannequin and agent harness beats general-purpose frameworks
Alongside the launch, Perplexity printed a research paper arguing that efficient native brokers require the mannequin and the agent harness — the scaffolding of prompts, instruments, and orchestration logic round the mannequin — to be designed collectively. The core perception: general-purpose harnesses assume a frontier mannequin that may take up huge contexts, navigate sprawling software surfaces, and plan over lengthy horizons. Small native fashions buckle beneath these calls for.
Perplexity discovered empirically that though fashions like Qwen 3.8 27B promote 260,000-token context home windows, they start to wrestle past 100,000 tokens. So the firm constructed a intentionally minimal harness: a succinct system immediate, a small set of core instruments, and capabilities that load and unload as on-demand “abilities” fairly than sitting completely in context. It transformed in style connectors like Gmail and GitHub from token-hungry MCP servers into compact command-line instruments, added self-verification hooks that monitor the well being of a job, and enforced always-on OS-level sandboxing. If the sandbox is unavailable, the harness disables itself fairly than operating instruments unprotected — a distinction with open-source harnesses that run instructions with the person’s full permissions by default.
The benchmark outcomes Perplexity reviews are putting, although they arrive from the firm’s personal evaluations. On its inner Local Knowledge Work Bench — 53 duties spanning deep analysis, monetary evaluation, and doc creation, which Perplexity says it plans to open-source — Laptop operating Qwen 3.8 27B on a DGX Spark scored 82.6%, versus 77.6% for the open-source Pi harness and 74.0% for Hermes operating the an identical mannequin.
Perplexity’s post-trained PPLX 27B pushed the rating to 85.4%. The gaps widen dramatically on more durable duties: on BrowseComp, an online analysis benchmark, Laptop hit 66.7% accuracy versus 50.2% for Pi and 43.9% for Hermes, whereas utilizing 51% much less wall time and 70% fewer tokens than Pi. On multimodal doc understanding, Laptop scored 65.1% towards Hermes’ 34.6% and Pi’s 13.9%.
The token economics driving AI brokers from the cloud to native {hardware}
The strategic logic behind the launch turns into clear when you think about how AI workloads have modified. Chat was bursty — a query, a solution, achieved. Brokers are completely different.
“With brokers, you need these brokers all the time on if you happen to can. You need the brokers to actually eat as many tokens as they will,” Nader stated. “What we’re seeing is an insatiable demand for tokens, and that is one thing that makes native AI so nice. As you noticed by all these demos, you have been not metered by the token. You have been not paying for the token. So it is actually killer for brokers.”
This reframes the worth proposition of native {hardware}. An agent that runs for hours reviewing paperwork, verifying its personal work, and iterating on analyses would rack up substantial API payments in the cloud. On a tool the person already owns, the marginal value of these tokens approaches zero. Perplexity’s paper makes the enterprise model of this argument explicitly: as brokers scale throughout particular person workflows and whole organizations, token expenditure and information motion “develop into more and more tough to govern.” Native-first execution addresses each directly — spend, as a result of inference is free, and privateness, as a result of delicate tokens by no means go away the machine boundary.
Maybe the most commercially fascinating outcome considerations the hybrid center floor. On Terminal Bench 2.1, a difficult coding benchmark, the totally native Qwen mannequin scored 59.6% at primarily zero marginal value. Letting it escalate to a Claude Opus 5 “advisor” in the cloud raised the rating to 73.0% at an estimated $0.415 per job. Working the frontier mannequin alone scored 82.4% at $0.65 per job. Escalation, in different phrases, recovered roughly three-fifths of the hole to frontier efficiency at about two-thirds of the value — and the person decides when that commerce is value making. Earlier than any advisor name, the harness runs a PII classifier over the outgoing context and reveals the person precisely what would depart the machine. The distant mannequin returns textual content steerage solely; it by no means touches native information or instruments.
The place Transportable Laptop matches towards Ollama and the DIY native AI stack
Jason Hiner of The Deep View pressed the firms on how Portable Computer relates to current native inference instruments like Ollama. Nate’s reply drew a transparent line: the instruments resolve completely different layers of the downside.
“The vast majority of the effort right here has been at the agent harness stage,” he stated, noting that the system makes use of vLLM to host mannequin inference beneath, with a complicated mode for customers who need to plug in their very own inference endpoint. “We have closely post-trained each the Qwen and Nemotron fashions that we’re working with so as to actually get the absolute best outcomes… Our focus has been on actually honing the entire stack, high to backside, of the mannequin inference and the harness collectively.”
Nader put it extra colorfully. “Simply getting inference operating actually shortly on a Spark — there is a easy path. You should use Ollama. You may get that arrange. However then, as you begin to do extra sophisticated, extra agentic issues, then out of the blue you want extra perf. You begin taking a look at completely different fashions. You begin taking a look at completely different harnesses, and it is form of like the ocean. The deeper you go, the deeper it will get.”
The appliance-like pitch appeared to land with at the very least one attendee. Ben, who described struggling to arrange his personal DGX Spark regardless of being an engineer — “this expertise sucks, we’ve to repair it” — stated the product seems like the unlock “wanted for folks to actually really feel and perceive what agentic means, and also you want the proper UX to make it occur.” Nvidia additionally emphasised that the {hardware} scales: connecting two Sparks over shared reminiscence runs frontier-class open fashions like DeepSeek’s newest, and 4 can run GLM 5.2 or Nemotron Extremely. “I’ve even seen eight Sparks get related,” Nader stated.
What the deepening Nvidia-Perplexity alliance means for each firms
The launch extends a partnership that has been constructing for greater than a yr. In June 2025, Nvidia and Perplexity announced a collaboration to carry sovereign AI fashions to European publishers and telecoms, a part of CEO Jensen Huang’s continent-hopping marketing campaign to persuade governments that, as the Related Press reported from VivaTech in Paris, “every country needs a national intelligence infrastructure.” The sovereign AI pitch — that information “belongs to your folks, your nation, your tradition,” in Huang’s phrases — is philosophically the similar argument Transportable Laptop makes at the scale of a single desk: intelligence you management, operating on {hardware} you personal.
There is a self-interested logic for each firms. Perplexity, which has raised capital at steadily escalating valuations whereas dealing with authorized strain from publishers over its content material practices — together with a lawsuit filed by The New York Times in December 2025 and an earlier public dispute with Forbes — will get a product whose economics do not rely on metering each token, and a differentiated wedge into privacy-sensitive enterprises in regulation, healthcare, and finance. Nvidia will get a killer app for DGX Spark, a tool that, by the admission of attendees at Monday’s briefing, has been simpler to purchase than to use. When one reporter requested whether or not a Spark would possibly ship with Transportable Laptop and a Nemotron mannequin preinstalled, Nader demurred with out ruling it out: “That might be cool… the objective is simply ensuring that it is a tremendous easy expertise for each person.”
Questions stay. Perplexity’s most spectacular numbers come from its personal inner benchmark, and the firm acknowledges that compact fashions nonetheless path the frontier meaningfully on arduous reasoning duties — advisor escalation “narrows however does not totally shut the hole.” The launch is Linux-only for now, the 24GB VRAM flooring excludes the overwhelming majority of shopper PCs, and Apple silicon — dwelling to a few of the most enthusiastic native AI tinkerers — is conspicuously absent from the roadmap. “We’re very targeted proper now on Nvidia {hardware},” Nate stated when requested.
However the route of journey is unmistakable. Perplexity’s researchers describe the launch as a part of “a broader shift wherein more and more succesful brokers transfer from distant infrastructure to particular person and native units,” and each firms are betting that advances in chips and open fashions will hold increasing what a field on a desk can do. Throughout Monday’s demos, the most telling element wasn’t a benchmark rating — it was that credit score counter in the nook of the display screen, sitting immobile at zero whereas the agent churned by a folder of tax paperwork. For 2 years, the AI trade has measured its ambitions in gigawatts and tokens per greenback. Transportable Laptop proposes a special meter, one which by no means runs.
Disclaimer: This article is sourced from external platforms. OverBeta has not independently verified the information. Readers are advised to verify details before relying on them.