ASUS GX10 & GB10 local models

Two ASUS Ascent GX10 nodes powered by the NVIDIA GB10 Grace Blackwell Superchip — 128 GB unified memory per node, ~1 PFLOP FP4 each, and a private Ollama stack for agent-grade inference on-premises.

NVIDIA GB10 Grace Blackwell 128 GB LPDDR5x NVLink-C2C ConnectX-7 @ 200 Gbps DGX OS Wi-Fi 7

NVIDIA GB10 Grace Blackwell Superchip

The GB10 integrates a Grace Arm CPU and Blackwell GPU on one coherent memory domain. CPU and GPU share 128 GB LPDDR5x via NVLink-C2C — about 5× the bandwidth of PCIe 5.0 — so large models stay resident without constant host/device copies.

Compute

ArchitectureNVIDIA Grace Blackwell
CPU20-core Arm v9.2-A (10× Cortex-X925 + 10× Cortex-A725)
GPUBlackwell generation (integrated)
CUDA coresBlackwell generation
Tensor cores5th generation (FP4 / FP8 / FP16)
RT cores4th generation
AI performanceUp to ~1 PFLOP (FP4 tensor)
SOC TDP140 W (240 W system peak)

Memory & interconnect

System memory128 GB LPDDR5x coherent unified
Memory interface256-bit
Memory bandwidthUp to 273 GB/s (8533 MT/s)
CPU ↔ GPU linkNVLink-C2C (coherent)
Model capacityFine-tune / infer up to ~200B params (single node)
Dual-node scale2 PFLOP FP4 · 256 GB unified · via ConnectX-7 QSFP
Cluster scale4+ nodes over network switch (ASUS-supported)

ASUS Ascent GX10 hardware

A 1.48 kg desktop AI supercomputer — 150 × 150 × 51 mm — with enterprise I/O, MIL-STD-810H rugged design, and a full NVIDIA AI software stack preloaded on DGX OS.

Storage

  • 1 TB / 2 TB M.2 2242 NVMe PCIe 4.0 ×4
  • 4 TB M.2 2242 NVMe PCIe 5.0 ×4 option
  • Single SSD slot (factory sealed chassis)

Networking

  • 1× RJ-45 10 GbE LAN
  • 1× NVIDIA ConnectX-7 SmartNIC (200 Gbps QSFP)
  • Wi-Fi 7 (2×2) · Bluetooth 5.4 LE
  • Multi-node over QSFP / switch fabric

I/O & power

  • 3× USB 3.2 Gen 2×2 Type-C (20 Gbps, DP alt)
  • 1× USB-C PD in (EPR PD3.1, 180–240 W)
  • 1× HDMI 2.1a · Kensington lock
  • 240 W external adapter · 1× NVENC / 1× NVDEC

Full platform summary

Spec ASUS Ascent GX10
Model / SKU familyASUS Ascent GX10 (GB10 Grace Blackwell)
Dimensions (W × D × H)150 × 150 × 51 mm (5.91 × 5.91 × 2.01 in)
Weight1.48 kg (3.26 lb)
Operating systemNVIDIA DGX OS (Linux-based, factory configured)
AI software stackCUDA · cuDNN · TensorRT · PyTorch · TensorFlow · Jupyter · NVIDIA NIMs
Supported model familiesLlama 3.x · DeepSeek R1 · Qwen · Gemma · Nemotron · Meta / Google frameworks
CertificationsBSMI · CB · CE · FCC · UL · CCC · Wi-Fi · MIL-STD-810H
Typical use casesLocal LLM inference · agentic AI · fine-tuning · RAG · ComfyUI image pipelines

Home lab fleet

gx10-eed4 · primary dev · Ollama · Hannah + Nicole agents
gx10-46ab · secondary node · overflow inference
Private mesh · ComfyUI · Cursor agents · no public endpoints

Why GB10 on the desk

  • 128 GB unified memory holds 27B–70B models quantized on one node
  • FP4 tensor path for fast agent loops and tool calling
  • Data stays local — no cloud inference for family or business drafts
  • ConnectX-7 ready when a second GX10 doubles capacity to 256 GB

Local models & capabilities

Pulled from the live Ollama inventory on gx10-eed4. All workloads are business-appropriate and family-safe.

Qwen3.8 27B

27.3B params · Q4_K_M · ~17.7 GB on disk

Primary

General-purpose agent model with tool use, chain-of-thought reasoning, and vision input. Used for drafting, code assist, and multi-step Hermes agent tasks.

Completion Tools Thinking Vision
Context
262,144 tokens
Embedding dim
5,120
Best for
Agents · long docs · analysis

Gemma 4 31B

31.3B params · Q4_K_M · ~19.9 GB on disk

Active

Google's Gemma 4 line — strong instruction following and structured output. Good for summaries, email drafts, and lightweight tool-calling workflows.

Completion Tools Thinking
Format
GGUF
Family
gemma4
Best for
Summaries · writing · QA

Nemotron 3 Nano 30B

31.6B params · Q4_K_M · ~24.3 GB on disk · MoE

Active

NVIDIA Nemotron hybrid MoE — optimized for efficient inference on Grace Blackwell. Excels at long-context retrieval, reasoning chains, and tool orchestration with lower active parameter cost.

Completion Tools Thinking
Context
1,048,576 tokens
Embedding dim
2,688
Best for
RAG · mega-context · research

Previously hosted & rotation pool

Llama 3.x (8B–70B)

General inference · chat · instruction

DeepSeek R1 (distilled)

Multi-step reasoning · math · planning

Mistral / Mixtral 8×7B

MoE chat · retired from active slot

Phi-3 / Phi-4

Small fast models · edge tests

CodeLlama / StarCoder2

Code completion · refactor assist

nomic-embed-text

Embeddings · RAG indexing

SDXL / Flux (ComfyUI)

Image generation · Krea-2 Turbo pipelines

Llama 3.2 Vision

Image + text · document OCR tests

Qwen 2.5 Coder

Code agents · CI log analysis

Public hardware overview only — no Tailscale IPs, API keys, or internal endpoints. Live model list synced from Ollama on gx10-eed4.