Hakobi

Insight

NVIDIA GB10 for Local LLMs: Who It Fits (and Who It Doesn't)

Published by Hakobi · 19 September 2026

A GB10 system is a small desktop box that can hold a very large AI model in memory — up to roughly 200 billion parameters at 4-bit precision — and run it on your own premises, with no per-token cloud bill and no data leaving the building. The trade-off is speed: it's built for one developer or a small team, not for serving dozens of users at once. That's the short answer to why NVIDIA's GB10 machines — including the xFusion FusionXpark and the Dell Pro Max with GB10, both of which Hakobi supplies — are in such high demand, and who they actually suit.

What GB10 actually is

The NVIDIA GB10 Grace Blackwell Superchip pairs a 20-core Arm CPU with a Blackwell GPU, and both share a single pool of 128 GB of LPDDR5x memory. The FusionXpark and the Dell Pro Max with GB10 are two vendors' builds on that same platform, both running NVIDIA's DGX OS with the NVIDIA AI software stack pre-installed. xFusion lists its unit at 150 × 150 × 50.5 mm and 1.2 kg, with self-encrypting NVMe storage in 1, 2 or 4 TB options.

GB10 platformNVIDIA's stated figure
Memory128 GB unified LPDDR5x, shared by CPU and GPU
Memory bandwidth273 GB/s
Peak computeUp to 1 PFLOP at FP4 — a theoretical figure that assumes sparsity
InferenceModels up to about 200 billion parameters
Fine-tuningModels up to about 70 billion parameters
NetworkingConnectX-7 NIC at 200 Gbps; NVIDIA lists up to four linked systems
Power supply240 W

Whether a specific model supports linking multiple units varies by vendor build, so confirm that for the exact unit before planning around it.

Why demand is so high

  • Memory capacity, not just speed. A large model simply won't load on a typical desktop GPU with a few tens of GB of dedicated memory. A 128 GB unified pool means models that previously needed a rack server can be loaded on a desk.
  • Data stays on-site. Prompts, contracts, client records and source code never pass through a third-party API. That's a real advantage for regulated or client-confidential work — though on-premises hosting doesn't by itself make a deployment PDPA-compliant, any more than cloud hosting makes it non-compliant; access control, logging and backup still have to be built and tested.
  • No per-token bill. Once the hardware is in place, usage isn't metered, which changes how freely a team will use AI day to day.
  • Same stack as the data centre. Because it runs NVIDIA's own software environment, work built on a GB10 can move to larger NVIDIA infrastructure later without being rewritten.
  • Small and low-power. It runs from a 240 W power supply on an ordinary desk — no server room required.

Use cases

1. A private LLM server for a small team. Internal chat, and question-answering over company documents (retrieval-augmented generation) for accounting, advisory, legal or engineering teams who can't send client material to a public AI service.

2. Coding assistants and AI agents. Running a coding model locally keeps proprietary source code in-house and removes usage limits for developers who lean on it all day.

3. AI development and prototyping. Building and testing AI applications, and fine-tuning models up to about 70 billion parameters, before committing budget to data-centre GPUs.

4. Edge and vision workloads. xFusion lists edge applications such as predictive maintenance and medical image analysis. For sites with limited bandwidth — a common Sabah constraint — running inference on-site avoids pushing raw video or sensor data to a cloud service.

5. Research and training. A self-contained AI workstation for universities, training providers and R&D teams that need hands-on access to large models without a shared cluster.

What it's not good at

This is where GB10 coverage online tends to skip ahead, so it's worth being direct.

  • The 1 PFLOP headline is theoretical. NVIDIA's own footnote qualifies it as FP4 with sparsity. It isn't what you'll see on most workloads.
  • Generation speed is limited by memory bandwidth. At 273 GB/s, producing each new token is slower than the compute headline suggests. HotHardware's review of the Dell Pro Max with GB10 measured about 10.65 tokens per second on a 31B-parameter Gemma 4 model versus 21.9 on an Apple Mac Studio M4 Max, and described local LLM performance as sluggish for interactive use.
  • Model choice matters a lot. The vLLM team's guidance is that mixture-of-experts models in NVFP4 with roughly 10–15 billion active parameters are the best fit. They measured about 22.7–23.7 tokens per second on a 120B-parameter mixture-of-experts model — a recipe-specific result, not a universal ceiling — while dense models of similar size are far less comfortable.
  • It's not a high-concurrency server. The vLLM team advises keeping concurrent requests low; beyond about four simultaneous streams, the bandwidth cost can outweigh the batching gains. Serving a whole company's chatbot needs a proper GPU server, not this.
  • 128 GB is shared. The operating system, model weights and the working memory for long conversations all draw from the same pool.

Is a GB10 the right fit?

Five questions worth answering before choosing hardware:

  • How many people will use it at the same time — one developer, a team of five, or fifty?
  • What's the largest model you actually need, and can a quantized or mixture-of-experts version do the job?
  • Is data confidentiality the driver, or is cloud cost?
  • Is this for development and experimentation, or a production service other systems depend on?
  • Who will secure it, update it, and back up the models, datasets and document indexes that end up on it?

If the answers point to one person or a small team, private data, and experimentation or light internal use, GB10 is a strong fit. If they point to many concurrent users or a business-critical service, a rack-mounted GPU server is the right tier — xFusion's server range sits in that space.

What Hakobi does around it

Choosing between the FusionXpark and the Dell Pro Max with GB10 is largely a question of vendor ecosystem and support relationship — the underlying platform is the same. We supply both, and we handle the parts that determine whether a local AI box is actually safe to rely on: network placement and access control so the model endpoint isn't exposed, backup of model weights, datasets and document indexes, monitoring, and power protection. Tell us the workload — users, model size, and what data it will touch — and we'll recommend a configuration. Pricing and availability depend on the specific model and configuration, so request a quote and we'll scope it from there.

Sources: NVIDIA DGX Spark specifications, NVIDIA; FusionXpark product page, xFusion; Dell Pro Max GB10 review, HotHardware; vLLM on the DGX Spark, vLLM Blog.

Related services

More insights

Ready to talk to an engineer?

Tell us about your site, your network, or your compliance deadline. We'll respond within one business hour.