Serving is the layer that turns a static model file into a running service — accepting requests, generating answers, serving many users. On Apple Silicon, three tools cover most needs: Ollama (easiest, start with one command), LM Studio (a desktop app with MLX and llama.cpp under the hood) and MLX (Apple's machine-learning framework, used through mlx-lm). vLLM is the option when you deploy on NVIDIA GPU servers with many concurrent users. For its internal AI packages on M5 Mac Studio, Namtech serves models with continuous batching (vllm-metal or oMLX) so many people can ask at once. Combine them with quantization (e.g. GGUF Q4/Q5/Q8, the llama.cpp format) to save memory and an OpenAI-compatible API so internal apps can call in. This is Part 4/8 of the build-your-own internal AI series.
Quick summary
- What serving is: software that loads the model and exposes an endpoint to accept requests — without it, the model is just a file on disk.
- Ollama: easiest; install, then run a model with a single command — the default choice for an internal AI server on a Mac.
- LM Studio: desktop app for downloading, testing and serving models, with MLX and llama.cpp under the hood and OpenAI-compatible endpoints.
- MLX: Apple's framework for Apple silicon;
mlx-lmruns and fine-tunes models natively on the Mac. - vLLM: only relevant on NVIDIA GPU servers — optimized for throughput with many concurrent users (continuous batching).
- Quantization: Q4/Q5/Q8 trade memory for quality; Q4 is usually a good balance to start.
- OpenAI-compatible API: a standardized
/v1/chat/completionsendpoint so any internal app calls in as if calling OpenAI.
Choosing a serving tool: Ollama, LM Studio, MLX or vLLM?
There's no single "best" tool — each is optimized for a situation. You choose by hardware type, concurrent users, and how much you want to configure yourself. On Apple Silicon, the choice is between Ollama, LM Studio and MLX; vLLM only comes in if you run NVIDIA GPUs.
- Ollama — easiest, start with one command. Compact install, manages models like images, exposes a local API out of the box. It runs GGUF models (the llama.cpp format) and also offers MLX builds of some models for Apple Silicon (for example the
qwen3.5:9b-mlxtag). A good default for one machine serving a team. - LM Studio — desktop app, no command line needed. LM Studio downloads and runs models with MLX and llama.cpp under the hood, and its local server exposes OpenAI-compatible endpoints (
/v1/chat/completions,/v1/embeddings…). Handy for testing and comparing models before settling on one. - MLX — native to Apple silicon. mlx-lm is a Python package for generating text and fine-tuning models on Apple silicon with MLX, using models from the Hugging Face Hub. Its built-in HTTP server is similar to the OpenAI chat API, but its own docs say it is not recommended for production, so put a proper server layer in front for real use.
- vLLM — for NVIDIA GPU servers. Built to serve many concurrent requests with continuous batching and efficient memory management on GPU. Consider it only if your deployment runs on NVIDIA GPUs rather than Apple Silicon.
| Criterion | Ollama | LM Studio | MLX (mlx-lm) | vLLM |
|---|---|---|---|---|
| Ease of getting started | Easiest (one command) | Easy (desktop app) | Moderate (Python) | More configuration |
| Best-fit hardware | Apple Silicon / GPU | Apple Silicon / PC | Apple Silicon only | NVIDIA GPU |
| Model format | GGUF, plus MLX builds for some models | GGUF (llama.cpp) and MLX | MLX weights from Hugging Face | Mostly native weights |
| OpenAI-compatible API | Yes | Yes | Similar (basic server) | Yes |
| When to use | Default internal AI server on a Mac | Testing & comparing models | Apple-native runs & fine-tuning | Only on GPU servers under high load |
The pragmatic path on Apple Silicon: start with Ollama (or LM Studio if your team prefers an app) to prove value, and use MLX when you want Apple-native builds or light fine-tuning. vLLM only makes sense if you choose NVIDIA GPU servers. See the Hardware article to pick the matching machine, such as a Mac Studio.
Quantization — GGUF Q4/Q5/Q8
Quantization is a technique that reduces the precision of model weights (e.g., from 16-bit down to 4-bit) so the model uses less memory and often runs faster. The GGUF format (used by llama.cpp, Ollama and LM Studio) ships common quantization levels ready to go: Q4, Q5, Q8. MLX models use their own quantized formats (e.g. 4-bit), with the same memory-versus-quality trade-off.
- Q4: the heaviest compression of the three, saving the most memory — the popular level because it balances size and quality well and is usually a reasonable starting point.
- Q5: moderate compression, a slight quality bump over Q4, in exchange for a bit more memory.
- Q8: light compression, keeping quality closest to the original of the three, but using the most memory.
| Level | Compression | Memory | Quality (vs original) | When to use |
|---|---|---|---|---|
| Q4 | Heaviest | Least | Balances size & quality well | Popular, safe starting point |
| Q5 | Moderate | A bit more than Q4 | Slight bump over Q4 | When Q4 falls short on important tasks |
| Q8 | Light | Most | Closest to the original of the three | When you need the highest quality |
The general rule: heavier quantization → less memory, but output quality may gradually drop. It's a trade-off, not free. The size of the effect depends on the model and task, so the right approach is to test on your own data rather than trusting a fixed percentage. For most enterprises starting out, Q4 is a safe launch point; if quality falls short on important tasks, move up to Q5/Q8 or pick a larger model (see the Model selection article).
Context length & concurrency/batching
Two parameters have a big impact on memory and latency when serving many people:
- Context length: the maximum tokens the model handles in one turn (prompt + answer). A longer context lets you feed more documents, but uses more memory and can be slower. Set it just large enough for the task rather than maxing it out "to be safe."
- Concurrency / batching: to serve many people at once, the engine groups requests and processes them in parallel. On GPU servers, vLLM stands out for continuous batching — raising total throughput under load. The trade-off: more concurrent requests → more memory needed, and the latency of any single request can rise as the queue lengthens.
In short: longer context + more concurrent users = more memory needed, and you have to balance throughput (total served) against latency (delay each person feels). Measure on your real load to pick the right configuration instead of guessing.
OpenAI-compatible API
A major strength of today's serving ecosystem is that most tools expose an OpenAI-compatible API — the same /v1/chat/completions endpoint shape, the same request/response structure. As a result, any internal app (chat UI, backend, automation script) can call your internal AI exactly as if calling OpenAI, just by changing the base URL to your on-premise server.
The benefits: standardization — existing libraries, examples and tools all work; and easy switching — if you later change engines (say Ollama → LM Studio, or to vLLM on a GPU server) or models, the apps calling in barely need edits. It's also the foundation for the next step — RAG — and for the UI & integration step.
See also the companion posts: Internal AI system architecture diagram, Internal AI security system and Trending Pool — updating world knowledge.
Fastest start — Ollama, run a model with a single command:
# install Ollama, then run an open-source model
ollama run qwen3.5:9b # chat right in the terminal, 100% offline
Apple-native with MLX — mlx-lm includes a basic HTTP server similar to the OpenAI chat API (fine for testing, not for production per its docs):
# serve a model from Hugging Face with MLX
mlx_lm.server --model <hf-repo-or-path> # listens on localhost:8080
Only if you run NVIDIA GPU servers — vLLM exposes an OpenAI-compatible server:
# serve a model via vLLM (OpenAI-compatible server)
vllm serve <model> # opens the /v1/... endpoint on GPU
Internal apps call the OpenAI-compatible endpoint — just point the base URL at your on-premise machine:
# call a chat completion on the internal server (OpenAI-compatible)
curl http://localhost:11434/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"qwen3.5:9b","messages":[{"role":"user","content":"Hello"}]}'
The Namtech view
Namtech builds the serving layer for internal AI running 100% on-site on Apple Silicon — each package is a single Mac Studio with an M5 chip (AI Box: M5 Max 64GB, AI Pro: M5 Max 128GB, AI Enterprise: M5 Ultra 256GB) — using an engine with continuous batching (vllm-metal or oMLX) so one machine serves many people asking at once, then tuning quantization and context against the customer's real load. We always standardize on an OpenAI-compatible API so your internal apps aren't locked into a specific engine — swapping models or engines later stays painless. Specific speed numbers depend on hardware, model and load, so we measure directly on your system rather than promising generic figures.
Frequently asked questions
Should I start with Ollama, LM Studio, MLX or vLLM?
On Apple Silicon, most should start with Ollama because you install it and run a model with a single command — fast to prove value. LM Studio suits teams that prefer a desktop app, and MLX suits Apple-native builds and fine-tuning. vLLM is only relevant if you run NVIDIA GPU servers with many concurrent users. All of them offer an OpenAI-compatible (or similar) API, so the apps calling in barely need edits.
Does Q4 quantization hurt the model much?
Quantization is a trade-off: the heavier the compression, the more memory you save but the more quality may gradually drop. Q4 is popular because it balances well and is usually a reasonable starting point. The size of the effect depends on the model and task — test on your own data; if it falls short on important tasks, move up to Q5/Q8.
Can I run internal AI without a GPU?
Yes. On Apple Silicon, Ollama, LM Studio and MLX use the unified memory shared by CPU and GPU, so no discrete GPU is needed; llama.cpp also runs on plain CPU. The speed and the model size you can run depend on the hardware — see the Hardware article to size a configuration by number of users.
What does "OpenAI-compatible API" mean?
It means the serving engine exposes an endpoint with the same shape as OpenAI's API (e.g., /v1/chat/completions). So internal apps call your on-premise AI exactly as they'd call OpenAI, just by changing the base URL to your server — no need to rewrite the integration.
Want internal AI without starting from zero?
Namtech deploys private internal AI platforms — open-source models running 100% on your own infrastructure, data never leaving the organization.
Book a free consultationNote: This is a general guide, last updated 01/10/2026; tools and models change fast — verify the latest versions when you deploy.