NVIDIA DGX Spark 64GB puts AI compute and memory at the center of a compact desktop system. It targets developers, researchers, students, data scientists, and AI enthusiasts who want to build with local models instead of relying on cloud services. The system combines the NVIDIA GB10 Grace Blackwell Superchip with 64GB of coherent unified memory and NVIDIA’s AI software stack. It supports model fine-tuning, inference, and autonomous AI agents. Built-in ConnectX-7 networking lets users scale to multi-Spark clusters as workloads grow. The 64GB configuration is available exclusively through OEM partners Acer, ASUS, Dell, Gigabyte, HP, and MSI.

A conventional desktop can handle basic AI experiments. Things change when the work grows more ambitious. Running one model is manageable. Keeping a larger model, its context, and the rest of an application available at the same time is harder. DGX Spark approaches the problem differently. It treats local AI as the core purpose of the machine rather than an afterthought on a general-purpose PC.

Fine-Tuning Models on Your Own Data

Local AI becomes more useful once the model is no longer treated as a finished product. A developer building a coding assistant may not want a generic model that simply knows how to program. Fine-tuning an existing model against a particular codebase creates a different starting point. NVIDIA specifically calls out this workflow for DGX Spark. Developers can fine-tune an existing LLM using their own software source code and then run the resulting model locally for future development tasks.

The same setup supports computer-vision projects, local inference testing, and data-science work. Developers can move from training and experimentation to testing without immediately shifting the workload elsewhere. The hardware shortens the distance between an idea and a test. A developer can alter the model, run it locally, inspect the result, and try again without sending every iteration to an external inference service. Local computing gives the person building with the model direct control over the experiment.

Hardware Built Around AI Workloads

DGX Spark 64GB is designed around that workflow. The GB10 Grace Blackwell Superchip pairs a Blackwell GPU with a 20-core Arm CPU. NVLink-C2C creates a coherent memory architecture so the processor and GPU work from the same memory pool. NVIDIA rates the system for up to 1 petaFLOP of FP4 AI performance and 273GB/s of memory bandwidth.

Memory matters more with local models than the GPU headline numbers often suggest. The model needs room to run. Longer contexts, additional models, and the rest of an application need space too. A machine that handles one model comfortably can struggle when the workload starts acting like a real application. NVIDIA positions the 64GB configuration around newer open models including Qwen3.8-27B, Meta Muse Glimmer, and NVIDIA Nemotron 3.5 Lightning. These models deliver capabilities comparable to frontier models from only a few months earlier while fitting inside a 64GB memory footprint.

Using CX-7 networking and NVIDIA Sync Cluster Assistant, two 64GB systems can be connected to pool 128GB of memory and run bigger models and larger workloads. More memory gives room to keep richer context available, run supporting workloads alongside the model, and experiment with more complicated local applications before the hardware becomes the limiting factor.

NVIDIA has also reduced some of the usual setup friction. DGX Spark runs DGX OS and ships with the NVIDIA CUDA accelerated AI software stack, including PyTorch, Jupyter, and Ollama for prototyping, fine-tuning, and inference.

Running Autonomous Agents Locally

An autonomous agent changes the experiment. The model is no longer expected to produce one answer and stop. A coding agent might inspect a project, identify a problem, make a change, test the result, and continue from what it has learned. A research workflow could involve several specialist models, tool calls, and a long trail of context that needs to stay available as the task progresses.

DGX Spark is built for those workloads. Multiple models and agents can share the same large memory pool. NVIDIA describes the system as a platform for long-context reasoning, multi-agent pipelines, and local inference. NVIDIA OpenShell provides policy-based privacy and security controls for agent workloads.

An enthusiast does not need a sprawling multi-agent system to use the hardware. A first project might be an agent that works through a personal dataset, checks a software project, organizes information, or handles a repetitive sequence of tasks. Local compute makes those experiments easier to iterate because the person building the workflow controls the models, the data, and the environment.

Cloud services still matter when a project needs enormous amounts of compute. Local hardware changes the equation for the experimentation that happens before a project reaches that scale. Every additional inference request no longer adds another usage cost. NVIDIA positions DGX Spark as a way to keep models, prompts, data, and inference on-device without per-token cloud inference fees.

Scaling Beyond a Single System

A project that starts with one machine does not have to stay there. NVIDIA built networking into DGX Spark so additional systems can join the same local AI environment as workloads grow.

Two DGX Spark 64GB systems can be connected to create a 128GB memory pool and up to 2 PFLOPS of total AI compute. NVIDIA states the pair can provide up to 1.7x the performance of a DGX Spark 128GB system in the configuration described in its materials. ConnectX-7 supplies the high-speed interconnect. NVIDIA Sync can discover the systems and help configure the cluster through its Cluster Assistant.

Larger setups extend the idea further. NVIDIA describes clusters of up to four DGX Spark systems. These create substantially larger shared memory pools for models and workloads that exceed the capacity of a single unit. Technical material also outlines distributed inference and fine-tuning workloads designed to scale across multiple systems.

NVIDIA Sync can stay in the background. The software discovers connected systems, configures the networking, routes workloads, and monitors system health. That lets the developer focus on the AI project instead of turning cluster setup into a separate task.

Compact Form Factor and Availability

The physical design stands out. DGX Spark measures 150mm by 150mm by 50.5mm and weighs 1.2kg. The hardware is still designed for workloads involving large models, local inference, and autonomous agents. NVIDIA describes the system as power-efficient enough to run from a standard wall outlet. That makes it practical for workloads that need to remain active rather than starting only when someone sits down at a workstation.

DGX Spark 128GB remains available through NVIDIA and all OEM partners for users who need more memory. The 64GB model serves as a practical starting point for someone who wants to experiment with local AI using the latest generation of small open-source models.

Putting this much AI capability on a desk changes the starting point of an experiment. A developer does not have to begin by deciding which cloud service to use or how much computing time a project deserves. The model, the data, and the tools can all sit within reach. That makes it easier to keep refining an idea until it becomes something useful.

DGX Spark 64GB is designed for that kind of work. Its combination of Grace Blackwell compute, 64GB of unified memory, and NVIDIA’s local AI software stack gives developers and AI enthusiasts room to work with larger models, fine-tune them, build agents, and test the results locally. When a project eventually needs more memory or compute, additional DGX Spark systems can extend the same setup rather than forcing a completely different workflow.

NVIDIA is unveiling the DGX Spark 64GB on October 2, 2026. The system becomes available from NVIDIA OEM partners on October 23.