Internal AI

Serving: install & optimize model speed

Serving internal AI models on-premise: Ollama, vLLM, llama.cpp and an OpenAI-compatible API

Serving is the layer that turns a static model file into a running service — accepting requests, generating answers, serving many users. On Apple Silicon, three tools cover most needs: Ollama (easiest, start with one command), LM Studio (a desktop app with MLX and llama.cpp under the hood) and MLX (Apple's machine-learning framework, used through mlx-lm). vLLM is the option when you deploy on NVIDIA GPU servers with many concurrent users. For its internal AI packages on M5 Mac Studio, Namtech serves models with continuous batching (vllm-metal or oMLX) so many people can ask at once. Combine them with quantization (e.g. GGUF Q4/Q5/Q8, the llama.cpp format) to save memory and an OpenAI-compatible API so internal apps can call in. This is Part 4/8 of the build-your-own internal AI series.

Quick summary

  • What serving is: software that loads the model and exposes an endpoint to accept requests — without it, the model is just a file on disk.
  • Ollama: easiest; install, then run a model with a single command — the default choice for an internal AI server on a Mac.
  • LM Studio: desktop app for downloading, testing and serving models, with MLX and llama.cpp under the hood and OpenAI-compatible endpoints.
  • MLX: Apple's framework for Apple silicon; mlx-lm runs and fine-tunes models natively on the Mac.
  • vLLM: only relevant on NVIDIA GPU servers — optimized for throughput with many concurrent users (continuous batching).
  • Quantization: Q4/Q5/Q8 trade memory for quality; Q4 is usually a good balance to start.
  • OpenAI-compatible API: a standardized /v1/chat/completions endpoint so any internal app calls in as if calling OpenAI.

Choosing a serving tool: Ollama, LM Studio, MLX or vLLM?

There's no single "best" tool — each is optimized for a situation. You choose by hardware type, concurrent users, and how much you want to configure yourself. On Apple Silicon, the choice is between Ollama, LM Studio and MLX; vLLM only comes in if you run NVIDIA GPUs.

  • Ollama — easiest, start with one command. Compact install, manages models like images, exposes a local API out of the box. It runs GGUF models (the llama.cpp format) and also offers MLX builds of some models for Apple Silicon (for example the qwen3.5:9b-mlx tag). A good default for one machine serving a team.
  • LM Studio — desktop app, no command line needed. LM Studio downloads and runs models with MLX and llama.cpp under the hood, and its local server exposes OpenAI-compatible endpoints (/v1/chat/completions, /v1/embeddings…). Handy for testing and comparing models before settling on one.
  • MLX — native to Apple silicon. mlx-lm is a Python package for generating text and fine-tuning models on Apple silicon with MLX, using models from the Hugging Face Hub. Its built-in HTTP server is similar to the OpenAI chat API, but its own docs say it is not recommended for production, so put a proper server layer in front for real use.
  • vLLM — for NVIDIA GPU servers. Built to serve many concurrent requests with continuous batching and efficient memory management on GPU. Consider it only if your deployment runs on NVIDIA GPUs rather than Apple Silicon.
CriterionOllamaLM StudioMLX (mlx-lm)vLLM
Ease of getting startedEasiest (one command)Easy (desktop app)Moderate (Python)More configuration
Best-fit hardwareApple Silicon / GPUApple Silicon / PCApple Silicon onlyNVIDIA GPU
Model formatGGUF, plus MLX builds for some modelsGGUF (llama.cpp) and MLXMLX weights from Hugging FaceMostly native weights
OpenAI-compatible APIYesYesSimilar (basic server)Yes
When to useDefault internal AI server on a MacTesting & comparing modelsApple-native runs & fine-tuningOnly on GPU servers under high load

The pragmatic path on Apple Silicon: start with Ollama (or LM Studio if your team prefers an app) to prove value, and use MLX when you want Apple-native builds or light fine-tuning. vLLM only makes sense if you choose NVIDIA GPU servers. See the Hardware article to pick the matching machine, such as a Mac Studio.

Quantization — GGUF Q4/Q5/Q8

Quantization is a technique that reduces the precision of model weights (e.g., from 16-bit down to 4-bit) so the model uses less memory and often runs faster. The GGUF format (used by llama.cpp, Ollama and LM Studio) ships common quantization levels ready to go: Q4, Q5, Q8. MLX models use their own quantized formats (e.g. 4-bit), with the same memory-versus-quality trade-off.

  • Q4: the heaviest compression of the three, saving the most memory — the popular level because it balances size and quality well and is usually a reasonable starting point.
  • Q5: moderate compression, a slight quality bump over Q4, in exchange for a bit more memory.
  • Q8: light compression, keeping quality closest to the original of the three, but using the most memory.
Table — Comparing GGUF quantization levels
LevelCompressionMemoryQuality (vs original)When to use
Q4HeaviestLeastBalances size & quality wellPopular, safe starting point
Q5ModerateA bit more than Q4Slight bump over Q4When Q4 falls short on important tasks
Q8LightMostClosest to the original of the threeWhen you need the highest quality

The general rule: heavier quantization → less memory, but output quality may gradually drop. It's a trade-off, not free. The size of the effect depends on the model and task, so the right approach is to test on your own data rather than trusting a fixed percentage. For most enterprises starting out, Q4 is a safe launch point; if quality falls short on important tasks, move up to Q5/Q8 or pick a larger model (see the Model selection article).

Your infrastructure — on-premise · data never leaves the org ClientInternal app APIOpenAI-compatible Serving engineBatching ModelQuantized Response
The flow of a single inference request, entirely within your infrastructure boundary: Client → API (OpenAI-compatible) → Serving engine (batching) → Model (quantized) → response. Diagram: Namtech.

Context length & concurrency/batching

Two parameters have a big impact on memory and latency when serving many people:

  • Context length: the maximum tokens the model handles in one turn (prompt + answer). A longer context lets you feed more documents, but uses more memory and can be slower. Set it just large enough for the task rather than maxing it out "to be safe."
  • Concurrency / batching: to serve many people at once, the engine groups requests and processes them in parallel. On GPU servers, vLLM stands out for continuous batching — raising total throughput under load. The trade-off: more concurrent requests → more memory needed, and the latency of any single request can rise as the queue lengthens.

In short: longer context + more concurrent users = more memory needed, and you have to balance throughput (total served) against latency (delay each person feels). Measure on your real load to pick the right configuration instead of guessing.

OpenAI-compatible API

A major strength of today's serving ecosystem is that most tools expose an OpenAI-compatible API — the same /v1/chat/completions endpoint shape, the same request/response structure. As a result, any internal app (chat UI, backend, automation script) can call your internal AI exactly as if calling OpenAI, just by changing the base URL to your on-premise server.

The benefits: standardization — existing libraries, examples and tools all work; and easy switching — if you later change engines (say Ollama → LM Studio, or to vLLM on a GPU server) or models, the apps calling in barely need edits. It's also the foundation for the next step — RAG — and for the UI & integration step.

See also the companion posts: Internal AI system architecture diagram, Internal AI security system and Trending Pool — updating world knowledge.

For the IT team

Fastest start — Ollama, run a model with a single command:

# install Ollama, then run an open-source model
ollama run qwen3.5:9b # chat right in the terminal, 100% offline

Apple-native with MLX — mlx-lm includes a basic HTTP server similar to the OpenAI chat API (fine for testing, not for production per its docs):

# serve a model from Hugging Face with MLX
mlx_lm.server --model <hf-repo-or-path> # listens on localhost:8080

Only if you run NVIDIA GPU servers — vLLM exposes an OpenAI-compatible server:

# serve a model via vLLM (OpenAI-compatible server)
vllm serve <model> # opens the /v1/... endpoint on GPU

Internal apps call the OpenAI-compatible endpoint — just point the base URL at your on-premise machine:

# call a chat completion on the internal server (OpenAI-compatible)
curl http://localhost:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"qwen3.5:9b","messages":[{"role":"user","content":"Hello"}]}'

The Namtech view

Namtech builds the serving layer for internal AI running 100% on-site on Apple Silicon — each package is a single Mac Studio with an M5 chip (AI Box: M5 Max 64GB, AI Pro: M5 Max 128GB, AI Enterprise: M5 Ultra 256GB) — using an engine with continuous batching (vllm-metal or oMLX) so one machine serves many people asking at once, then tuning quantization and context against the customer's real load. We always standardize on an OpenAI-compatible API so your internal apps aren't locked into a specific engine — swapping models or engines later stays painless. Specific speed numbers depend on hardware, model and load, so we measure directly on your system rather than promising generic figures.

Frequently asked questions

Should I start with Ollama, LM Studio, MLX or vLLM?

On Apple Silicon, most should start with Ollama because you install it and run a model with a single command — fast to prove value. LM Studio suits teams that prefer a desktop app, and MLX suits Apple-native builds and fine-tuning. vLLM is only relevant if you run NVIDIA GPU servers with many concurrent users. All of them offer an OpenAI-compatible (or similar) API, so the apps calling in barely need edits.

Does Q4 quantization hurt the model much?

Quantization is a trade-off: the heavier the compression, the more memory you save but the more quality may gradually drop. Q4 is popular because it balances well and is usually a reasonable starting point. The size of the effect depends on the model and task — test on your own data; if it falls short on important tasks, move up to Q5/Q8.

Can I run internal AI without a GPU?

Yes. On Apple Silicon, Ollama, LM Studio and MLX use the unified memory shared by CPU and GPU, so no discrete GPU is needed; llama.cpp also runs on plain CPU. The speed and the model size you can run depend on the hardware — see the Hardware article to size a configuration by number of users.

What does "OpenAI-compatible API" mean?

It means the serving engine exposes an endpoint with the same shape as OpenAI's API (e.g., /v1/chat/completions). So internal apps call your on-premise AI exactly as they'd call OpenAI, just by changing the base URL to your server — no need to rewrite the integration.

Want internal AI without starting from zero?

Namtech deploys private internal AI platforms — open-source models running 100% on your own infrastructure, data never leaving the organization.

Book a free consultation

Note: This is a general guide, last updated 01/10/2026; tools and models change fast — verify the latest versions when you deploy.

Get started

Start with a free assessment

To define the right package and detailed scope, Namtech offers a short, no-cost assessment.

We reply within 1 business day. No spam, we never share your info.