Posting measured observations after installing Perplexity Computer v26.9.0 (arm64) on a
DGX Spark (GB10, 128 GB unified), since I didn’t find community notes on its memory footprint and what leaves the machine. Everything captured below is measured with ss, tshark, nvidia-smi, and the app bundle itself. Summary below pulled from Claude Code CLI.
What’s good
- The local model server is a standard OpenAI-compatible endpoint at
http://127.0.0.1:45793/v1``(llama-server style), bound to loopback (not0.0.0.0). So it’s composable — you can point your own tooling at it — and not exposed to the LAN. - Cloud escalation is gated per-action. Every step that would leave the box (web search,
adviser call, GitHub PR, etc.) throws an explicit Allow/Deny with a plain-language “what leaves.” - On-device inference runs fully on Spark; no per-token cost for local steps.
Memory reality (biggest disappointment)
- Install pulls PPLX 27B (Qwen-27B post-trained, 4-bit, ~27.6 GB download). PPLX is
selectable as the orchestrator; Nemotron 3.5 Lightning (~19 GB) is “coming soon.” - A 27B 4-bit model is ~15–20 GB of weights. But the loaded vLLM engine reserved ~82 GB
(nvidia-smi: 83,984 MiB on one engine core) on the 128 GB unified box. - That’s vLLM’s default
gpu-memory-utilization ≈ 0.9KV-cache preallocation, not model size. On unified memory (CPU+GPU shared) it pushed system RAM to ~121 GB, into swap, and the GPU allocator hitNVRM ... NV_ERR_NO_MEMORY— any other GPU app (I tried launching Firefox) crashes while the model is loaded. There is no user-facing knob to lower the reservation (user-settings.jsononly picks the model). - “Close window” ≠ quit. Closing left ~12 Electron processes + the model resident. Right-click
the local model → Stop to release the ~82 GB; the app minimizes to background rather than exit.
Takeaway:
Perplexity Computer behaves like the DGX Spark is its dedicated inference server and expects to own the box — may be great on a headless node, rough as a desktop app if you are on the Spark through NVIDIA OS. Plan for one-model-owns-the-machine reserving memory for orchestration.
Telemetry (heavier than “local-first” implies)
Static scan of /opt/Perplexity shows an embedded analytics stack: Sentry, Datadog, Segment, Heap, Amplitude, New Relic. Observed on the wire (tshark, SNI/DNS):
- Persistent HTTPS to Datadog intakes (
browser-intake-datadoghq.com,http-intake.logs.us5.datadoghq.com). - ~195 Datadog “span” batches generated locally in a single ~2–3 hr session (
~/.config/Perplexity/spans/). - At idle (no task, no apps, browser closed) the Perplexity Computer app holds ~8 persistent connections to its backend (Cloudflare-fronted) + Google/GCP. Toggling Remote Access off dropped ~4 of them (the GCP relay).
- The crash reporter is mandatory — the app FATALs on launch if
chrome_crashpad_handlercan’t spawn.
None of this is required for local inference to run — it’s vendor-side observability/product
analytics. Blocking it has no functional downside to on-device work but it’s hard-set as mandatory out telemetry for Perplexity Computer to work.
One caveat worth knowing: the Privacy Gate has a gap
The pre-escalation Privacy Gate checks plain-text and code files only. Per its own disclaimer, PDFs, Office docs, images, audio, and video are sent without a PII check — i.e., exactly the formats sensitive documents usually are. For confidential-doc workflows, keep them local and don’t approve escalation.
Practical tips
- Reclaim memory: right-click model → Stop; don’t rely on closing the window.
- See what it talks to:
ss -tunp | grep -i perplexand
sudo tshark -i any -Y 'tls.handshake.type==1' -T fields -e tls.handshake.extensions_server_name. - Silence the behavioral telemetry (null-route both IPv4 and IPv6 in
/etc/hosts, since the app connects over v6):browser-intake-datadoghq.com,http-intake.logs.us5.datadoghq.com,
o<id>.ingest.us.sentry.io,api.segment.io,api.heapanalytics.com,api.amplitude.com,bam.nr-data.net. Leave*.perplexity.ai(that’s the app’s functional backend). - Heads-up: blocking the intakes makes the local span batches pile up (they can’t upload), so
sweep~/.config/Perplexity/spans/on a cron.
Initial Verdict
Promising, but unless the app behavior changes, budget for a dedicated node — the model appears to claim the whole box — and go in knowing the “local-first” label covers inference, not the telemetry plane. Unless you choose to share tracking, terminate the behavioral analytics; it maps usage for the vendor, not performance for you.
Side Note:
Perplexity Computer offers access to your Enterprise/Pro subscription so your local activity can be accessed online if you toggle remote on, you can also allow for web search and escalation; so you can choose local-only without remote access, web search or escalation, or local-accessible where you can access your local files while allowing cloud escalation to Perplexity frontier models, web search/ deep research, remote access from your Perplexity account.
These modes govern the app’s intentional egress. The observability/telemetry plane and the idle backend connection persist regardless — so treat the mode selector as convenience, not a guarantee. True local-only comes from enforcing it at the network (see the /etc/hosts block above), not from the app dropdown.
Will share more impressions after running a few local orchestration tasks.
Tested: DGX Spark GB10 / 128 GB / DGX OS (aarch64), Perplexity Computer 26.9.0, Perplexity Enterprise Max.