Internal AI

On-premise hardware for internal AI: how to choose

Choosing on-premise hardware for internal AI: Apple Silicon and GPU

Hardware for internal AI is decided by four factors: model size (parameter count), number of concurrent users, context length, and the speed you expect. The pivotal factor is memory — RAM or VRAM must hold the model weights plus context. For most enterprises, Apple Silicon (Mac Studio or Mac mini, unified memory, low power draw) is a tidy starting point; NVIDIA GPUs fit when you need very high throughput or training. This is Part 2/8 in the build-your-own internal AI series.

Quick summary

  • What decides it: model size × concurrent users × context length × desired speed — all of which reduce to a memory requirement.
  • Apple Silicon vs GPU: Apple Silicon has unified memory, so it runs large models on shared RAM with low power, small footprint and quiet operation; NVIDIA GPUs give high throughput and suit training but cost more in power, heat and space.
  • Memory rule: RAM/VRAM ≈ (parameter count × bytes per quantization) + a portion for context — this is a rule of thumb, not an exact figure.
  • Sizing by scale: AI Box (1× M5 Max 64GB, a small team) → AI Pro (1× M5 Max 128GB, several departments) → AI Enterprise (1× M5 Ultra 256GB, whole enterprise), each a single M5-chip Mac Studio. The number of users served is determined by a load test on real documents, not fixed per package.
  • Namtech: deploys on low-power Mac Studio M5 Max / M5 Ultra machines, scaling up as needs grow.

What decides your hardware needs?

Before debating "which machine to buy", answer four questions — they drive everything else:

  • Model size (parameter count): a bigger model (7B, 14B, 32B, 70B…) is smarter but consumes more memory and runs slower. This is the single biggest variable.
  • Concurrent users: one person asking occasionally is very different from 30 people typing at once. Many concurrent users need higher throughput (often a GPU or several machines).
  • Context length: letting the AI read long documents, long conversations or many RAG passages consumes extra memory for the "context" (KV cache) — the longer the context, the more memory.
  • Desired speed: is "fast enough to read along" acceptable, or do you need near-instant replies? Higher speed expectations demand more powerful hardware or a smaller model.

The key point: all four factors reduce to memory and throughput. Once you pick a model size and user count, the rest (machine type, RAM/VRAM capacity) follows fairly naturally. See how the pieces fit together in the internal AI system architecture diagram.

Apple Silicon or GPU?

This is the biggest hardware decision. Both run open-source models; they differ in memory architecture, power draw and the situations they suit.

CriterionApple Silicon (Mac Studio / Mac mini)NVIDIA GPU
MemoryUnified memory — CPU & GPU share it, so a single machine can load a large model if configured with high RAMDedicated VRAM per card — powerful, but per-card capacity is limited; large models must span multiple cards
Power & heatLow power, cool, quiet — fine in a normal officeHigh power draw & heat, usually needs a server room/cooling
Multi-user throughputGood for small–medium teams; scale by adding machinesVery high, suited to serving many concurrent users
Heavy training / fine-tuneLight workloads only; not its strengthIts strength — the mature CUDA ecosystem for training
Size & installationCompact, plug in and goBulkier, needs matching power & cooling

For most enterprise tasks (internal assistant, document Q&A, drafting, summarizing) and moderate user counts, Namtech chooses Apple Silicon — a Mac Studio M5 Max or M5 Ultra per package: unified memory lets a single machine load a fairly large model, low power draw means it sits in a normal office, and all three packages fit in a single machine. For smaller pilots, a Mac mini M6 or M5 Pro is an option — see the Mac mini vs Mac Studio comparison. When needs lean toward very high throughput or heavy training, an NVIDIA GPU is the more sensible choice.

Sizing by scale

Namtech packages hardware into three tiers, each one M5-chip Mac Studio. Prices are current selling prices including VAT (sold and invoiced by Decorp), taken from the product pages; models are those Namtech uses per package (updated 01/10/2026). The table does not fix a user count per package — the number of users you can serve depends on model size, context length and real load, so Namtech runs a load test on your own documents before committing.

TierMachineUnified memoryStorageModels Namtech uses
AI Box1× Mac Studio M5 Max (111,150,000₫)64GBSSD 512GBQwen3.6-35B-A3B or Gemma 4 26B-A4B
AI Pro1× Mac Studio M5 Max (161,310,000₫)128GBSSD 512GBQwen3.8-27B or Gemma 4 31B, alongside Qwen3.6-35B-A3B; SSO
AI Enterprise1× Mac Studio M5 Ultra (307,800,000₫)256GBSSD 1TBQwen3.5-122B-A10B or GLM-5.3-Flash

A pragmatic principle: start at the lowest tier that solves your first clear problem, measure the real impact, then expand — rather than over-buying up front. Details of the packages are on the pricing page, and every Mac Studio configuration with its estimated model size is on the Mac Studio page.

Scale up gradually — three packages, one Mac Studio each 1AI BoxMac Studio M5 Max64GB · 1 machineQwen3.6-35B-A3B / Gemma 4 26B-A4B 2AI ProMac Studio M5 Max128GB · 1 machineQwen3.8-27B / Gemma 4 31B2 models side by side · SSO 3AI EnterpriseMac Studio M5 Ultra256GB · 1 machineQwen3.5-122B-A10B / GLM-5.3-Flash Users & load increase →
Three internal AI packages, each one M5-chip Mac Studio; the number of users served is determined by a load test on real documents. Diagram: Namtech.

The memory rule of thumb

This is the most important part of choosing a configuration. Memory (RAM on Apple Silicon, VRAM on GPU) must hold:

  • Model weights: roughly parameter count × bytes per parameter. Bytes per parameter depend on quantization — compressing the weights to use less memory. At the common Q4 level (about half a byte per parameter), a ~7B model needs only about a few GB; larger models (14B, 32B, 70B) need proportionally more. For reference, Namtech's Mac Studio page estimates roughly ~43B parameters for 36GB, ~76B for 64GB and ~115B for 96GB of unified memory.
  • Context (KV cache): the memory for context — the longer the context or the more concurrent users, the more this grows.

Add the two together, then leave a safety margin for the operating system and load variation. This is a rule of thumb for a quick estimate, not an exact figure for every model — real numbers vary by architecture and quantization level, so check the specific model page on Hugging Face. Exactly which model size and quantization to pick is covered in the Model selection article.

Storage & networking

Beyond memory and processing, three often-overlooked items directly affect the experience:

  • SSD: models can be many GB and must load quickly into memory at startup. The AI Box and AI Pro machines ship with a 512GB SSD and AI Enterprise with 1TB; plan an SSD large enough to hold multiple models plus the vector database for RAG (indexed internal documents).
  • Internal network: users call the AI server over the LAN. A stable internal network keeps responses smooth and keeps everything within the on-premise boundary — no data pushed outside.
  • UPS (battery backup): so models and services don't shut down abruptly during a power cut, avoiding data corruption and disruption.

How to lock the data flow inside the network and control access is covered in the Internal AI security system article.

For the IT team

How to check machine resources and the models you have, plus a quick memory-estimation rule:

  • List installed models: ollama list — shows the name and size of each model on the machine.
  • Check resources: free RAM, GPU/VRAM (on Apple Silicon it's unified memory, so total RAM is the number to watch).
  • Estimate memory: ≈ (parameter count × bytes per quantization) + a portion for context, then leave a safety margin. For exact numbers, check the model page on Hugging Face.
# list the models on this machine
ollama list
# rough estimate: 7B at Q4 ~ a few GB; larger models need more
# pull & run a small model to measure real memory use
ollama run qwen3.5:9b

On a Mac you can also run models with MLX or LM Studio; Activity Monitor shows the memory each one really uses.

The Namtech view

Namtech deploys private internal AI platforms on Apple Silicon — one Mac Studio M5 Max or M5 Ultra per package: unified memory lets a machine load a fairly large model, low power draw means it sits in a normal office, and you move up a tier as load grows — instead of over-investing up front. Not sure between a Mac mini and a Mac Studio? See the comparison page. The philosophy is start with just enough, scale gradually: pick a configuration for your first clear problem, measure the real impact, then upgrade as needed. The next step is to choose the open-source model that fits the hardware you've picked.

Frequently asked questions

Is Apple Silicon or GPU better for internal AI?

There's no absolute answer. Apple Silicon (Mac Studio / Mac mini) has unified memory, so it runs large models on shared RAM with low power, a small footprint and quiet operation — a fit for most enterprise tasks at moderate scale. NVIDIA GPUs give high throughput and excel at training, but cost more in power, heat and space. Namtech uses Apple Silicon (Mac Studio) for most deployments.

How much memory does a 7B model need?

As a rule of thumb, a ~7B model at Q4 quantization needs about a few GB for weights, plus a portion for context. That's a quick estimate — real numbers vary by architecture and quantization level, so check the specific model page on Hugging Face.

What configuration do I need for the whole company?

It depends on model size, concurrent users and context length. Namtech's packages: AI Box (1× Mac Studio M5 Max 64GB) for a small team or one department, AI Pro (1× M5 Max 128GB) for several departments, AI Enterprise (1× M5 Ultra 256GB) for the whole company. The number of users served is not fixed per package; Namtech runs a load test on your own documents before committing.

Do I need a dedicated server room?

With low-power Apple Silicon machines, usually not — the machines are compact, cool and quiet, and can sit in a normal office (a UPS is advisable). High-wattage GPUs need matching power and cooling, in which case a server room makes sense.

Want internal AI without starting from zero?

Namtech deploys private internal AI platforms — open-source models running 100% on your own infrastructure, data never leaving the organization.

Book a free consultation

Note: The memory figures in this article are rules of thumb; package configurations and models last updated 01/10/2026; hardware and models change fast — verify the specific configuration against your real needs when you deploy.

References
Get started

Start with a free assessment

To define the right package and detailed scope, Namtech offers a short, no-cost assessment.

We reply within 1 business day. No spam, we never share your info.