โก NVIDIA RTX Spark for Local AI (2026): What It Actually Runs
128GB unified memory, a 20-core Grace CPU, a Blackwell GPU on one chip. What RTX Spark can actually run locally, what it costs, and who should wait to buy.
Derek Holt ยท Local AI & Hardware Writer
ยท 4 min read

RTX Spark packs a 20-core Grace CPU, a Blackwell RTX GPU, and up to 128GB of unified memory onto one chip, and NVIDIA says the flagship configuration hits one petaflop of AI compute while still gaming at roughly 100fps at 1440p. That's the headline. The number that actually matters for a local-AI buyer is 128GB shared memory on a chip that isn't priced yet and isn't shipping yet.
I run an RTX 4080 for everything on this site, 16GB of VRAM, and I hit its ceiling constantly on anything past a 13B-class model at full precision. A chip that claims 8x that pool changes the math, if it ships at a price a hobbyist can justify.
By the numbers
| RTX Spark (N1X, flagship) | My RTX 4080 rig | |
|---|---|---|
| Memory | 128GB unified (shared CPU+GPU) | 16GB dedicated VRAM |
| CPU | 20-core Grace (Arm) | Consumer x86, separate from GPU |
| AI compute | 1 petaflop claimed | Not petaflop-class |
| CUDA cores | 6,144 | 9,728 |
| Streaming multiprocessors | 48 | 76 |
| Power envelope | 45โ80W | 320W TDP |
| Claimed model ceiling | 120B params locally (200B+ with two linked units) | ~30B at usable quant, realistically |
| Price | Unconfirmed, analyst estimates $1,799โ$3,000 | ~$1,000 street (used market varies) |
| Availability | Fall 2026, OEM date TBD | Shipping now |
The power number is the one I keep coming back to. A 45โ80W envelope running a 120B model is not a rounding error against a 320W discrete card โ it's a different category of machine, closer to a laptop than a rig with a case fan problem.
The correction most coverage skips
RTX Spark and DGX Spark get conflated constantly, and it matters because they answer different questions. DGX Spark ships now, runs the GB10 superchip under DGX OS (Linux), targets models up to 200B parameters, and comes with NVIDIA AI Enterprise software. It's a developer box, priced and positioned like one. RTX Spark is a different silicon design called N1X, boots Windows on Arm, is scoped to 120B-class models, and ships inside consumer laptops and compact desktops from OEMs who haven't announced pricing.
Same memory philosophy. Different chip, different OS, different buyer. If an article tells you RTX Spark specs while showing a DGX Spark box, it hasn't checked which product it's describing.
What 128GB unified actually buys a local-AI user
Unified memory means the CPU and GPU draw from the same pool instead of a discrete card's dedicated VRAM plus separate system RAM. That's the same architectural bet Apple made with the M-series, and it's why RTX Spark gets compared to a "MacBook moment" for Windows.
Practically: a 70B-parameter model that needs roughly 40โ48GB at a usable quant currently means either a multi-GPU rig or accepting slow CPU offload on a single 24GB card. On paper, RTX Spark's 128GB pool clears that with room to spare, and NVIDIA's own claim of 120B locally implies headroom past 70B-class models entirely, without touching a server.
That's the promise. Nobody outside NVIDIA's own demos has independently benchmarked real-world tokens-per-second on retail hardware yet, because retail hardware doesn't exist yet.
What I'd actually wait to see before buying
Three numbers aren't public: real price, real memory bandwidth, and real sustained tokens-per-second on a 70B+ model outside a controlled demo. Bandwidth is the one I care about most โ unified memory pools are only as fast as the bus feeding them, and a wide-but-slow pool can lose to a narrower, faster discrete card on actual inference speed, not just capacity. I've been burned by spec-sheet VRAM numbers before; a memory pool you can't feed fast enough doesn't help.
The $1,799โ$3,000 estimated range is wide enough to mean completely different products. At the low end, RTX Spark undercuts a used RTX 3090 rig on capability-per-dollar for anyone who values a silent, low-power box. At the high end, it's competing with a used enterprise card, and the calculus flips.
Who should wait, who should watch
If you're running local LLMs today on a 12โ24GB card and hitting real ceilings on model size, RTX Spark is worth watching, not buying yet โ there's no price, no ship date beyond "fall," and no independent bandwidth numbers. If you're comparing it against building a discrete rig right now, our best GPU for local AI guide and the RTX 5090 vs 4090 breakdown cover what you can buy and run today, with real, tested numbers instead of a keynote slide.
For most hobbyists, that's still the better bet this year. RTX Spark reads like the direction the platform is heading, not something to build a fall budget around yet.
How we picked
Specs (memory, CPU core count, CUDA core count, SM count, power envelope, petaflop claim) are from NVIDIA's own Computex and GTC Taipei 2026 announcements and confirmed OEM partner statements. The RTX Spark vs DGX Spark distinction is cross-checked against NVIDIA's developer forums and multiple independent silicon breakdowns. Pricing is explicitly unconfirmed. Every number above the official range is an analyst estimate, flagged as such, and should be treated as a placeholder until an OEM posts a real SKU.
My own RTX 4080 numbers are measured on my rig, not a spec sheet.
Frequently asked questions
โธIs RTX Spark the same as DGX Spark?
No, and this is the mistake half the coverage makes. They share a similar idea โ an Arm CPU and a Blackwell GPU on one package with up to 128GB of unified memory โ but DGX Spark runs on the GB10 superchip under Linux-based DGX OS and targets AI development up to 200B-parameter models, while RTX Spark runs on a separate chip called N1X, boots Windows on Arm, and is positioned for 120B-class local models plus RTX gaming and creator work. Two chips, two operating systems, two product lines.
โธHow much does RTX Spark cost?
Unconfirmed as of this writing. NVIDIA has not opened pre-orders or published official pricing. Analyst estimates circulating range from roughly $1,799 to $3,000 depending on configuration, but treat those as guesses, not quotes, until an OEM posts a real price.
โธWhen can I actually buy one?
NVIDIA says RTX Spark PCs arrive fall 2026, led by ASUS and MSI, with Acer, GIGABYTE, HP, and Microsoft following. No OEM has announced a specific on-sale date. If you need local AI compute today, this is a wait-and-see purchase, not a next-week one.
โธWhat's the actual VRAM-equivalent number that matters?
128GB of unified memory is the flagship configuration โ shared between CPU and GPU, which is a different architecture than a discrete card's dedicated VRAM. NVIDIA's own claim is that it runs 120-billion-parameter models locally, and two linked units can reportedly push past 200B โ territory that used to mean a server rack, not a desk.
The 5 best AI video finds, every week
New models, tested prompts, and what actually worked in our production โ one short email a week. No spam, unsubscribe anytime.
Written by Derek Holt
Local AI & Hardware Writer
Runs the site's local-inference rig and benchmarks every GPU, quant, and speed-stack claim on it personally before it goes in a guide. Will not shut up about VRAM bandwidth.
Explore these topics
Every guide, comparison and prompt library we have on each.
Keep learning
GPUs & Hardware ยท Local LLMs
ComparisonsRTX 5070 Ti vs RTX 3090 for Local AI: 16GB New or 24GB Used?
The 5070 Ti's 16GB warranty versus a used 3090's 24GB ceiling. Bandwidth, tokens/sec, power and price compared โ plus the one question that settles it.




