Using Local Models
Agency can run LLM calls on models stored on your own machine. Nothing leaves your computer, you need no API key, and the calls cost nothing. Local models are slower than hosted ones and less capable at the same size, so they shine for development, testing, offline work, and private data.
Local inference uses llama.cpp under the hood. Models are single .gguf files, mostly downloaded from Hugging Face.
Setup
Install the local provider once:
npm i -g smoltalk-llama-cppAgency finds the package automatically, whether you installed it globally or in your project. You can browse the catalog without it; any command that needs it will tell you to run this install.
Browse the catalog
agency local listThe first line shows the directory models are downloaded to. Below it, you get the full catalog of curated models, with a checkmark next to the ones you have already downloaded:
Models directory: /Users/you/.agency-agent/models
NAME PARAMS SIZE CONTEXT LICENSE
✓ smollm2-135m 135M 0.11 GB 8K apache-2.0
qwen3.5-2b 2B 1.28 GB 128K apache-2.0
gpt-oss-20b 20B 12.00 GB 128K apache-2.0
...For longer descriptions of each model, run agency local alias list.
Picking a model: the SIZE column is roughly what the model takes on disk and in memory, so it is the main thing to match against your hardware. Start small. smollm2-135m downloads in seconds and is good for checking that everything works. qwen3.5-2b is a reasonable first model for real tasks on a laptop.
Download a model
agency local downloadWith no argument, this opens a picker so you can choose from the catalog. You can also name a model directly:
agency local download qwen3.5-2bThe value can be a curated name, one of your aliases, a Hugging Face URI, or a path to a .gguf file you already have:
agency local download hf:Qwen/Qwen2.5-7B-Instruct-GGUF:Q4_K_MDownloads of curated models are verified against pinned SHA-256 hashes. A file that fails verification is set aside and never loaded.
You do not have to download ahead of time. Everything that runs a local model downloads it first if it is missing. Pre-downloading just moves the wait to a moment you choose.
See where models live
The first line of agency local list names the models directory. By default it is ~/.agency-agent/models. To change it, set the AGENCY_MODELS_DIR environment variable, or set client.modelsDir in agency.json:
{
"client": {
"modelsDir": "/data/agency-models"
}
}To free up disk space, delete a model's files with -f. Without -f the command only removes the alias and tells you where the files are.
agency local remove qwen3.5-2b -fRefresh the catalog
The catalog of curated models updates over time. Pull the latest list without upgrading agency:
agency local refreshNew and updated entries land in your agency.json as aliases. Any alias you added yourself is never overwritten. The local command reference covers the details, including pointing refresh at your own catalog.
Run the agent on a local model
agency agent --local qwen3.5-2bThis downloads the model if needed and points every LLM call at it, including the deep subagents. A fully local session needs no hosted API key at all. Run agency agent --local with no value to pick from the catalog interactively.
Run a program on a local model
The --local flag on agency run pins the whole run to a local model:
agency run --local qwen3.5-2b hello.agencyThe value accepts the same forms as agency local download: a curated name, an alias, an hf: URI, or a .gguf path. The download and verification happen before your program starts, so you see the progress in your terminal. --local and --model are mutually exclusive.
This composes well with run budgets. A run that must not spend money can say so explicitly:
agency run --max-cost 0 --local qwen3.5-2b hello.agencyUse local models in code
For finer control, resolve a model in code and pass it to llm yourself. The registerLocalModel function downloads the model if needed and returns its local path:
import { registerLocalModel } from "std::agency/local"
node main() {
const model = registerLocalModel("qwen3.5-2b")
const answer = llm("What is the capital of France?", {
model: model,
provider: "llama-cpp",
})
print(answer)
}Structured output and tool calls work the same way they do on hosted models:
type Capital = {
city: string
population: number
}
node main() {
const model = registerLocalModel("qwen3.5-2b")
const answer: Capital = llm("What is the capital of France?", {
model: model,
provider: "llama-cpp",
})
print(answer.city)
}This is useful when one program mixes models, for example routing cheap classification to a local model and hard reasoning to a hosted one. std::agency/local also exports the rest of the CLI's capabilities as functions: downloadModel, listModelNames, listDownloadedModels, aliasModel, and more. See the std::agency/local reference.
To make local the default for a whole project instead, set it in agency.json:
{
"client": {
"defaultProvider": "llama-cpp",
"defaultModel": "/Users/you/.agency-agent/models/my-model.gguf"
}
}Name your own models with aliases
An alias gives a short name to any model URI, so your team can share one agency.json and write llm(..., { model: registerLocalModel("my7b"), ... }) everywhere:
agency local alias add my7b hf:Qwen/Qwen2.5-7B-Instruct-GGUF:Q4_K_M
agency local alias list
agency local alias remove my7bAliases work everywhere a model value is accepted: agency local download, agency run --local, agency agent --local, and registerLocalModel.
Run MLX models on a Mac
MLX is Apple's array framework for Apple Silicon. Agency can run models through it as well as through llama.cpp. MLX models are often faster than GGUF models on a Mac, and the community publishes them at sizes llama.cpp builds rarely reach.
An MLX model is not one file. It is a directory of .safetensors weights next to a config.json, the layout Hugging Face repos use. Agency does not run it in its own process. A Python program called mlx_lm.server loads the model, and Agency sends it requests. Agency ships a small script around it that adds structured output, using the llguidance library, so a typed llm() call gets a reply that fits its type.
Set up Python once:
python3.12 -m venv ~/.agency-agent/mlx-env
~/.agency-agent/mlx-env/bin/pip install mlx-lm==0.31.3 llguidance==1.8.0Agency looks for that environment by default. Agency never installs Python for you. To use a different Python, pass --python, or set client.mlx.python in agency.json.
Download a model
agency local download mlx:mlx-community/Qwen3.8-27B-4bitAn mlx: URI names a Hugging Face repo. The download fetches the repo in parallel and verifies each file against the hash Hugging Face publishes. Interrupt it and run the same command again, and it picks up where it stopped.
Browse what is available with agency local list. MLX entries show mlx in the BACKEND column, and their names end in -mlx when a GGUF entry of the same model exists.
Start the server
agency local serveWith no model named, this shows you the MLX and diffusers models you have downloaded, each with its kind (chat, embedding, speech, or image), and you pick the ones to serve. Name them yourself to skip the picker:
agency local serve mlx:mlx-community/Qwen3.8-27B-4bitThe command runs in the foreground and prints the address it is listening on. Leave it running in its own terminal. Ctrl-C stops it and every model it loaded.
You can serve several models at once. Agency starts one mlx_lm.server for each of them and puts a single port in front. Each model stays in memory for as long as the command runs, so watch the total against the memory your Mac has. Agency warns you when the models add up to more than that, and starts them anyway.
A server uses memory beyond its weights while it works, and it cuts short a reply that goes in circles. Both are covered under What is different about a local model below, along with the flags that tune them.
Run against the server
agency run --local mlx:mlx-community/Qwen3.8-27B-4bit hello.agency
agency agent --local mlx:mlx-community/Qwen3.8-27B-4bitThese work the same way they do for a GGUF model, with one difference. Agency does not download or load anything here. The server has to be running already, and it has to be serving the model you name.
A run sends the model name to the server, so the two must agree. agency local resolve <name> prints the name Agency will send.
Use a model you already have
agency local serve /Volumes/models/hf/hub/models--mlx-community--Qwen3.8-27B-4bitAny directory holding config.json and .safetensors files works, so a model another tool downloaded needs no copying. A Hugging Face cache folder works too. Agency reads its refs/main to find the revision you last pulled.
agency local --model-dir /Volumes/models/hf/hub list
agency local --model-dir /Volumes/models/hf/hub serve--model-dir points the whole local command at another directory for one run. The picker then offers everything in that directory. Agency reads a Hugging Face cache but never writes to one, so agency local remove -f refuses to delete a model in one.
Watch what the server is doing
The server prints one line per request:
POST /v1/chat/completions mlx-community/Qwen3.8-27B-4bit 200 2.6s 16→129 tokThat line gives you the endpoint, the model, the status, the time the request took, and the tokens in and out. Add --verbose to see the whole request and the whole reply as well:
agency local serve mlx:mlx-community/Qwen3.8-27B-4bit --verboseShorter names
agency local alias add coder mlx:mlx-community/Qwen3-Coder-Next-4bit
agency local serve coder
agency run --local coder hello.agencyAn alias works everywhere a model value is accepted, the same as it does for a GGUF model.
If you already have the model, you can also drop the prefix and use the repo id on its own:
agency local serve mlx-community/Qwen3-Coder-Next-4bitThe mlx: prefix is the spelling that always works. You need it for a model you have not downloaded yet, because that is what tells Agency where to fetch it from.
Differences from GGUF models
The table at the end of What is different about a local model sets the two backends side by side. One difference is not in it: the agent's memory feature needs an embedding model, and a chat server does not serve one unless you name one. Name it when you start the server, and name it for the agent too. serve knows an embedding model when it sees one, so no flag is needed:
agency local serve qwen3.5-27b-mlx qwen3-embedding-4b-mlx
agency agent --local qwen3.5-27b-mlx --model embedding=mlx/qwen3-embedding-4b-mlxIn agency.json, the same thing is embeddings: { model, provider: "mlx" } under the memory settings.
Generate images on a Mac
Four open image models run on this machine:
| Model | Good for | 1024×1024 on an M5 Ultra | Download |
|---|---|---|---|
z-image-turbo | Fast, photorealistic images | about 8 s | 32.8 GB |
chroma1-hd | Detailed, cinematic pictures | about 90 s | 27.5 GB |
qwen-image-2512 | Legible text inside the image: signs, posters, diagrams | not measured yet | 57.7 GB |
flux2-klein-4b | Macs with less memory; 4 steps per image | not measured yet | 16.0 GB |
z-image-turbo and chroma1-hd have no content filter in their weights; flux2-klein-4b is safety fine-tuned. They run on Hugging Face's diffusers library rather than MLX, so they need torch and diffusers in the same Python:
~/.agency-agent/mlx-env/bin/pip install torch==2.14.0 diffusers==0.40.0 transformers==5.17.0 accelerate==1.15.0 sentencepiece==0.2.2 protobuf==7.36.2You can serve images without installing any MLX packages; serve checks only for what the models you name need.
Download a model and serve it. serve reads what kind of model it is from the download, so it starts the image server without being told:
agency local download z-image-turbo
agency local serve z-image-turboThen call it from Agency code:
import { generateImageLocal } from "std::image"
node main() {
const r = generateImageLocal("a lighthouse in a storm", "z-image-turbo", seed: 7)
if (isFailure(r)) { print("failed: ${r.error}"); return }
writeBinary("lighthouse.png", r.value.base64)
print("made with seed ${r.value.seed}")
}The same prompt and seed make the same image. Leave seed out and the server picks one, and the result says which. steps and guidance default to each model's own settings. z-image-turbo and flux2-klein-4b take no guidance and no negative prompt, and say so if you pass one.
Image models use more memory while they generate than their size on disk: chroma1-hd is 27.5 GB on disk and peaks at 36 GB.
Your own style with a LoRA adapter
A LoRA adapter is a small file that teaches an image model a style or a character from a few dozen example pictures. SDXL models take them, and the illustration finetunes NoobAI-XL and Illustrious are the ones most adapters are trained for. Serve one by its repo, and name the folder your adapters are in:
agency local download diffusers:Laxhar/noobai-XL-1.1
agency local serve diffusers:Laxhar/noobai-XL-1.1{ "client": { "adaptersDir": "./adapters" } }Then name the adapter in the call, by its file name without .safetensors. A request that leaves lora out gets the plain model:
const r = generateImageLocal("sketch, a cat on a chair", "diffusers:Laxhar/noobai-XL-1.1", lora: "sketch", loraScale: 0.9)loraScale is how strongly the adapter is applied: 1 is as trained, less is subtler, and up to 2 is allowed. The server loads an adapter the first time it is asked for, so a file you drop into the folder after training works at once, and when you train it again under the same name, the next request uses the new weights. Adapters are .safetensors files only, and a request can only pick a file from the folder you configured, never name a path itself.
Pose a character with a ControlNet
A ControlNet holds the image to a drawing you give it, so a stick figure becomes the pose. Name a folder for them, download one, and pass a drawing with the call:
{ "client": { "controlnetsDir": "./controlnets" } }agency local download controlnet-scribble-sdxlconst r = generateImageLocal("pen and ink, zxq_girl, surprised", "diffusers:Laxhar/noobai-XL-1.1",
lora: "zxq", controlnet: "controlnet-scribble-sdxl", controlImage: "./poses/jump.png")The drawing is read on this machine under std::readImage, the one thing a local generation asks approval for. controlScale is how strongly the drawing constrains the image, 1 by default. controlnet-openpose-sdxl takes a rendered pose skeleton instead of a scribble.
The scribble ControlNet reads white lines on a black background. A drawing made with dark lines on white paper needs invertControlImage: true:
const r = generateImageLocal("a dancer mid-leap", "diffusers:Laxhar/noobai-XL-1.1",
controlnet: "controlnet-scribble-sdxl", controlImage: "./poses/pen-sketch.png", invertControlImage: true)The drawing is scaled to fit size with its shape kept, and centered on black. A 4:3 drawing in a square image gets black bands above and below. It is never stretched.
Look at images on a Mac
Two vision models turn a picture into words and boxes, on this machine, with nothing uploaded. wd14-tagger describes an image as booru tags, the vocabulary illustration models such as NoobAI-XL take in a prompt. florence-2 finds the objects you name and returns a box for each, tags what it sees, and writes a caption. Download and serve them like any other model:
agency local download wd14-tagger florence-2
agency local serve wd14-tagger florence-2Then ask. Each function does one thing and returns a Result, so a loop over a folder is your own code, and cropImage from std::image turns a box into a new file:
import { detectObjects, tagImage } from "std::vision"
import { cropImage } from "std::image"
import { glob } from "std::shell"
node main() {
const pages = glob("comics/*.png") catch []
for (page in pages) {
const found = detectObjects(page, ["person", "desk", "chair"], "florence-2") catch []
for (hit in found) {
const out = "dataset/${hit.label}_${hit.id}.png"
cropImage(page, hit.box, out, pad: 0.05)
const tags = tagImage(out, "wd14-tagger") catch []
const names = map(tags) as t {
return t.tag
}
write("${out}.txt", names.join(", "))
}
}
}Run it with --approve std::glob --approve std::vision --approve std::cropImage --approve std::write and it labels a few hundred crops unattended. Every call raises an effect naming the file it reads or writes, so a policy can allow tagging under ./dataset and nothing else. Partial application narrows a tool before an agent gets it: detectObjects.partial(labels: ["person"], model: "florence-2") is a person-finder, and cropImage only ever writes a file that does not exist yet.
Find your own things
A detector finds things by the name of their kind. Ask for "mug" and it finds every mug, with no way to tell yours from the one next to it. Ask for "cat" on a pen-and-ink page and it may find nothing, because it learned mostly from photos. Two more models close both gaps without training anything:
owlv2-basefinds the things you name, likeflorence-2, andfindRegionsasks it to box every thing in a picture with no names at all.dinov2-baseturns a picture, or boxes in it, into embeddings: lists of numbers that say how each looks. Things that look alike get embeddings that are close together.
agency local download owlv2-base
agency local download dinov2-base
agency local serve owlv2-base dinov2-baseTo find your cats on a page: box everything, embed each box, and keep the boxes that are closer to crops of your cats than to crops of other things you draw.
import { findRegions, embedImage } from "std::vision"
import { cosineSimilarity } from "std::embedding"
// How alike `vector` is to the closest of `references`, where 1 is identical.
def closest(vector: number[], references: number[][]): number {
let best = 0
for (reference in references) {
const similarity = cosineSimilarity(vector, reference)
if (isSuccess(similarity) && similarity.value > best) {
best = similarity.value
}
}
return best
}
// One embedding per crop file.
def embedCrops(paths: string[]): number[][] {
let vectors: number[][] = []
for (path in paths) {
const embedded = embedImage(path, "dinov2-base")
if (isSuccess(embedded)) {
vectors.push(embedded.value[0])
}
}
return vectors
}
node main(page: string, catCrops: string[], otherCrops: string[]) {
const cats = embedCrops(catCrops)
const others = embedCrops(otherCrops)
const regions = findRegions(page, "owlv2-base")
if (isFailure(regions)) {
return regions.error
}
const boxes = map(regions.value) as region {
return region.box
}
const vectors = embedImage(page, "dinov2-base", boxes: boxes)
if (isFailure(vectors)) {
return vectors.error
}
for (vector, i in vectors.value) {
if (closest(vector, cats) > closest(vector, others)) {
print("Cat in region ${i}")
}
}
}Run it with --approve std::vision. Make the crops with findRegions and cropImage on a few pages, then sort the files into two folders by hand. Compare like with like: a tight crop and a box around the same thing compare well, but a whole photo of a mug on a desk mostly describes the desk. For your mug, use detectObjects(photo, ["mug"], "owlv2-base") for the boxes, since the detector does know what a mug is, and crops of other people's mugs as the other side.
How well this works has been measured on one photo so far. Two cats scored 0.55 against each other and at most 0.26 against a TV remote, so telling kinds of thing apart works. Telling your own mug from a lookalike in another photo has not been measured yet, and neither have drawings.
Chat with pictures on a Mac
~/.agency-agent/mlx-env/bin/python -m pip install mlx-vlm==0.7.0 mlx-audio==0.5.4
agency local serve --vlm qwen3.5-9b-mlxUse --vlm to ask a chat model questions about pictures. The initial supported architecture is Qwen3_5ForConditionalGeneration, tested with mlx-community/Qwen3.5-9B-4bit. Download the model first with agency local download qwen3.5-9b-mlx.
import { image } from "std::thread"
node main() {
const answer = llm(["What is in this picture?", image("photo.png")], {
provider: "mlx",
model: "mlx-community/Qwen3.5-9B-4bit"
})
return answer
}The image goes through the ordinary image-read approval. Set MLX_BASE_URL to the server's /v1 URL if it uses a port other than 8080. The model stays loaded between calls.
A typed call can ask for a specific value:
import { image } from "std::thread"
node main() {
const sides: number = llm(["How many sides does this shape have?", image("square.png")], {
provider: "mlx",
model: "mlx-community/Qwen3.5-9B-4bit"
})
return sides
}This is a chat model you ask questions in words. std::vision runs fixed tasks such as tagging, captioning, and detection.
The flag is required because these same weights also work as text models. Without it, Agency keeps using its existing text server. A model served with --vlm has no hedge or repeat limits, limitAnswers, or --draft. Thinking is off unless the call turns it on. Temperature and maxTokens still apply. Requests for unsupported reply limits fail with an error.
Both runtimes can share one serve command. Pinning mlx-audio in the install command preserves the version Agency's speech server requires. Use --python /path/to/env/bin/python if you keep a separate environment.
An alias keeps its resolved model identity: an mlx: URI uses the repo ID, and a directory uses its absolute path. GET /v1/models lists these names once the server is ready. Serving uses installed files with Hub networking and telemetry disabled. Missing files fail instead of downloading.
Chat bodies are limited to about 107 MB, including previous images resent in the conversation. Each image sent by the client is capped at 20 MiB. The local server accepts image bytes in data URIs and refuses remote URLs, file paths, audio parts, and alternate image part formats.
A reply you gave up on keeps running
With mlx-vlm 0.7.0, closing a streamed reply cancelled generation in the real-model check. Closing a non-streaming reply did not. The abandoned request finished before the next request generated its answer. A client timeout therefore does not guarantee that GPU work stops.
Agency starts mlx-vlm with its upstream log level set to CRITICAL. Upstream error messages can contain an invalid image's entire data URI. Agency's front door still logs request status and timings, and redacts image data from verbose request and response logs.
Use local models from TypeScript
A TypeScript program can call the same server with no Agency code. Import from agency-lang/local:
import { listModels, generateImage, tagImage } from "agency-lang/local";
import { writeFileSync } from "node:fs";
const generated = await generateImage({
model: "z-image-turbo",
prompt: "a lighthouse in a storm",
size: "1024x1024",
});
if (!generated.success) {
throw new Error(generated.error);
}
writeFileSync("lighthouse.png", generated.value.bytes);
const tags = await tagImage({ model: "wd14-tagger", image: generated.value.bytes });
if (tags.success) {
console.log(tags.value.map((tag) => tag.tag).join(", "));
}Start the server first, with both models: agency local serve z-image-turbo wd14-tagger.
generateImage takes the settings generateImageLocal takes, and detectObjects, tagImage, captionImage, embedImage, and findRegions take the settings of their std::vision namesakes. Each sends the request the stdlib function sends. Four things differ:
- No approvals. The stdlib functions ask before they read a file, because a model may have chosen the file. Your program chose it, so these functions read it directly.
- An image is a path or bytes. Pass
image: "page.png"orimage: someUint8Array. An image already in memory does not need to be written to a file first. - The server's address is an option. Pass
baseUrl: "http://127.0.0.1:8081/v1"to call a server on another port. Without it, the functions useclient.baseUrl.mlxinagency.json, thenMLX_BASE_URL, then port 8080. - A call can be cancelled. Pass
signalfrom anAbortController. Aborting it ends the call with the failureCancelled.
const cancelButton = new AbortController();
const pending = generateImage({ model, prompt, signal: cancelButton.signal });
cancelButton.abort();
const result = await pending; // { success: false, error: "generateImage failed: Cancelled" }listModels() returns the models downloaded to this machine. Each entry has the model's name, kind, family, directory, size, and the aliases that point at it:
const imageModels = listModels().filter((model) => model.kind === "image" && model.complete);Start the server from your program
serve starts the same server agency local serve does, and resolves once every model has loaded:
import { serve, generateImage } from "agency-lang/local";
const server = await serve(["z-image-turbo", "wd14-tagger"], {
log: (line) => console.log(line),
});
const generated = await generateImage({
baseUrl: server.url,
model: "z-image-turbo",
prompt: "a lighthouse in a storm",
});
await server.close();With no port, the server takes any free port, and server.url is its address. Pass it as baseUrl to each call. log receives what the command would print and what the model processes write. Without it, that output is discarded.
A model can be an object when it needs a setting the command gives with a flag:
await serve([
{ model: "qwen3.5-9b-mlx", vlm: true }, // --vlm
{ model: "/models/my-embedder", kind: "embedding" }, // --embedding
{ model: "qwen3.5-9b-mlx", draft: "qwen3.5-0.8b-mlx" }, // --draft
]);The server it returns can stop and start one model while the others keep running:
await server.unload("z-image-turbo"); // frees its memory
server.status(); // [{ model: "z-image-turbo", state: "stopped", ... }, ...]
await server.load("z-image-turbo"); // resolves when it is ready againA request for an unloaded model fails with a message that says to load it. server.failure is a promise that resolves with a message if a model process dies.
Any client can read the same status at GET /v1/agency/status on the server's address.
Serve more models than fit in memory
Mark a model lazy to load it on its first request, and to let the server stop it when another lazy model needs the memory:
const server = await serve([
"florence-2", // loads now, stays loaded
{ model: "z-image-turbo", lazy: true }, // loads on its first request
{ model: "flux2-klein-9b", lazy: true },
]);On the command line this is agency local serve florence-2 --lazy z-image-turbo --lazy flux2-klein-9b.
The first request for a lazy model waits for the load, and that wait counts against the caller's timeout. An image request with eight steps is given about three minutes, and a cold load of a large image model can take a good part of that. When a request has a deadline, warm the model first with server.load(model), which resolves when it is ready.
When a lazy model does not fit, the server stops the lazy model that has been idle longest and has no request running, and tries again. If every loaded model is busy or is not lazy, the request fails with a 503 that says so:
Not enough memory to load flux2-klein-9b (needs about 22 GB, 9 GB available).
Loaded now: z-image-turbo (busy), florence-2 (not lazy).A caller can wait and retry, or unload something itself.
Cancel a call, or everything on a model
To cancel your own call, pass a signal and abort it. The call fails with Cancelled, and the server closes its connection to the model process. For an image model that stops the work after the step in progress.
To stop whatever is running on a model, whoever started it:
await server.cancel("qwen3.5-9b");POST /v1/agency/cancel
{ "model": "qwen3.5-9b" }Every request in progress on that model fails with a 499. For most models the process is kept, because it stops on its own when the connection closes. A chat model served with vlm: true that was asked for a reply without streaming keeps generating after the connection closes, so for that one case the process is stopped and, unless the model is lazy, started again. That takes as long as loading the model.
Stop the server
await server.close();POST /v1/agency/shutdownclose refuses new requests, tells every model process to stop, waits up to five seconds for them to exit, forces any that remain, and then closes the port. The shutdown route does the same; under agency local serve the command then exits, and under serve() your program keeps running. If the server's own process is killed outright, each model process notices and exits on its own.
agency-lang/local is the supported way in. Files under agency-lang/stdlib-lib/ are internal to Agency, and their names change between releases.
There is no chat function here. A chat model served by agency local serve answers the OpenAI chat API at the same address, so any OpenAI client works.
What is different about a local model
A hosted provider makes a dozen small choices for you, and you never see them. A local model makes you see every one. This section lists the choices that catch people, what each looks like when it goes wrong, and what to do about it. Most apply to both backends. Where one applies to the MLX server alone, or to llama.cpp alone, the text says so.
The same prompt gives the same reply
Both local backends pick the single likeliest token at every step when nothing tells them otherwise. This is called greedy decoding. A greedy model writes the same reply for the same prompt every time, to the token, and takes the same time doing it. Every hosted provider samples instead, so its replies vary from call to call.
Agency makes a local model sample too when a call names no temperature, with the settings its model card asks for. Every catalog entry carries its card's numbers: Qwen3.5 asks for a temperature of 1.0 with a top-p of 0.95 and a top-k of 20, Gemma 4 for 1.0 with a top-k of 64, Mistral Small for 0.15, and so on. A model the catalog does not know gets 0.7 with a top-p of 0.95. The MLX server is sent all of it in the request. llama.cpp is sent the temperature and the cut-offs under node-llama-cpp's own names, and applies its own top-p of 0.95 and top-k of 40 to a model with no card of its own. Hosted providers sample at 1.0, but on top of samplers of their own; on the MLX server, 1.0 with nothing else is sampling from the whole distribution, which a small 4-bit model turns into noise. Ask for the repeatable behaviour when you want it:
import { setLlmOptions } from "std::llm"
setLlmOptions({ temperature: 0 })Two things follow from greedy decoding. Running a greedy model three times measures nothing the first run did not, so a benchmark at temperature 0 needs one trial. And a greedy model that starts going in circles cannot get out of them, because whatever led it to write "But wait" once leads it to write the same thing again.
Thinking spends your output budget
A thinking model writes a block of reasoning before its answer. Your program never sees that block, but it counts against maxTokens along with the answer. A small model can spend thousands of tokens thinking about a one-line question. When the budget runs out inside the thinking block, the call returns an empty reply with a stop reason of length, and a typed call fails with "reply did not fit the type".
Turn thinking off when the task does not need it:
import { setLlmOptions } from "std::llm"
setLlmOptions({ thinking: { enabled: false } })
const label: string = llm("Is this review positive or negative? ...")This asks the model's chat template to open no thinking block. On a classification or extraction task it saves most of the time a thinking model takes. Where thinking helps, give it a budget instead:
const answer: NumberAnswer = llm(problem, { thinking: { enabled: true, budgetTokens: 2048 } })After the budget, the thinking block is closed for the model and it has to answer. reasoningEffort works too, and means the same amount of thinking on a local model as it does on Gemini:
| effort | tokens of thinking |
|---|---|
| low | 2048 |
| medium | 8192 |
| high | 16384 |
Both backends honour all of this. On the MLX server the switch goes to the chat template and the budget to the server's watcher. On llama.cpp the switch goes to the chat wrapper, for the models whose wrapper has one, and the budget to llama.cpp itself. node-llama-cpp picks that wrapper by reading the model's template, and for a model it gets wrong, such as a fine-tune with a changed template, the switch reaches nothing. Name the wrapper yourself in agency.json with client.llamaCpp.chatWrapper, using node-llama-cpp's name for it (qwen, gemma4, harmony, chatML). Which models the switch reaches depends on their template: Qwen and Gemma 4 have one, gpt-oss takes a reasoning_effort instead (which reasoningEffort sends), and DeepSeek has none, so on DeepSeek off means a budget of zero and the block closes as soon as it opens. The budget is always held under maxTokens by enough to answer in.
Replies that go in circles
A model that is unsure can write "But wait, is that right? Let me reconsider." for thousands of tokens, or repeat one sentence until its budget runs out. With a hosted model this costs you money. With a local one it costs you the machine. A large model can take ten minutes to write a reply nobody wants, and every other call waits behind it.
The MLX server watches every reply for three signs of this:
- Thinking past its budget. The budget is half of
maxTokensunless the call setsbudgetTokens. - Twelve second thoughts in the last two thousand tokens. "But wait," "Wait," "Hmm," "Hold on," "Let me reconsider," and phrases like them. A loop says these every few lines; an honest long reply says them a dozen times over thousands of tokens, which is why the count runs over a window rather than the whole reply.
- The same sentence of six or more words, written three times in that window.
The last two watch the thinking only, unless you ask. An answer repeats itself for honest reasons: a refrain, a table with a repeated row, three similar functions in a file. Thinking rarely does. --limit-answers on agency local serve, or limit_answers: true on a request, watches answers too.
When the server sees a sign, it cuts the reply short in the way that leaves the most usable result. A thinking block is closed, so the answer can follow. A reply that must fit a type has the text field it is writing closed, so the rest of the type can still be filled in. A plain reply ends where it is. An answer cut short is reported to the client with a stop reason of length, the same as one that hit maxTokens; closing the thinking is not a cut, because the answer still comes. An Agency program does not see the stop reason, so a cut typed reply that still fits its type succeeds like any other. The server's log is where to look: it prints a line saying which limit tripped and what it did.
You can change the limits when you start the server. 0 turns one off:
agency local serve mlx:mlx-community/Qwen3.5-2B-4bit --hedge-limit 20 --repeat-limit 0A program can change them for its own calls, which matters when an answer repeats itself on purpose. replyLimits takes the same three settings; a field left out keeps the server's:
import { setLlmOptions } from "std::llm"
// A song's chorus repeats, and that is not a loop.
setLlmOptions({ replyLimits: { repeatLimit: 0 } })
const lyrics: string = llm("Write a song with a chorus that repeats after every verse.")
// A single careful call may hedge more than twelve times honestly.
const proof: string = llm(problem, { replyLimits: { hedgeLimit: 30 } })llama.cpp has no watcher. There, a reply that goes in circles runs to maxTokens or to the call's timeout, whichever comes first.
Every call has a cap and a clock
Two limits bound every call, and a local model hits both more often than a hosted one. The cap is maxTokens. llama.cpp caps a call at 16,384 tokens when you set none. The MLX server caps it at its --max-tokens, also 16,384 by default, and a call that asks for more gets that much. The clock is the runtime's per-call timeout of ten minutes. A call that hits it fails with "Request was aborted".
Set both lower for a local model, because a reply that goes in circles costs the whole cap and the whole clock:
setLlmOptions({ maxTokens: 8192, timeout: 300000 })timeout is in milliseconds. 8,192 tokens is room enough for any honest reply, a small thinking model's working included; a reply that needs more has almost always looped.
A reply you gave up on keeps running, unless something stops it
When a call times out, or you press Ctrl-C, your program moves on. The model does not know that. llama.cpp runs inside your process, so Agency stops it directly. The MLX server is another process behind a socket, and mlx_lm.server on its own only looks at that socket once the reply is finished. Agency's chat server looks every half second instead, and drops the reply at its next token. A prompt still being read is read to the end first, and on the path a draft model uses that means the whole prompt. The server's log shows it:
The client went away; stopping its reply.An abandoned reply that nothing stops runs to maxTokens, and every other request shares the GPU with it and runs two to three times slower. If a local model gets slower as a run goes on, check for this first.
Memory is the weights, plus everything the server remembers
A model's size on disk is the floor of what it needs, not the total. Each reply being generated holds attention state that grows with the prompt and the reply. On a 235B model that is about 190KB per token, so a 22,000-token prompt adds 4GB. The MLX server also keeps that state for the last ten prompts it saw, so a follow-up on the same conversation skips re-reading it. And the GPU keeps memory it has freed around for reuse.
None of this shrinks on its own. A long run on a large model grows until the GPU runs out of memory. The server then stops answering, and every later call waits out its timeout. The system log records the moment:
Execution of the command buffer was aborted due to an error during execution. Insufficient MemoryAgency limits the server to a sixteenth of the machine's memory for attention state, and the server drops its oldest prompts to stay under. Watch the Prompt Cache lines the server prints to see how much it holds.
The GPU does not get the whole machine. macOS gives it a working set below the total, and keeps the rest for everything else: on a 256GB Mac the GPU's share is about 223GB, so a 132GB model leaves about 90GB of that share for attention state and the prompt cache. The share can be raised, which takes memory from the rest of the machine until the next reboot:
sudo sysctl iogpu.wired_limit_mb=240000Do this only when a model almost fits and nothing else on the machine matters while it runs.
Which model fits the machine
A model's name says how its weights are stored. 4bit on an MLX model and Q4_K_M on a GGUF one both mean about four bits per weight instead of sixteen, which is a quarter of the size for a small loss in quality; 8bit and Q8_0 are half the size for almost none. The catalog picks 4-bit builds, and at 4 bits a model takes about 0.6GB per billion parameters: 27B is 16GB, 31B is 18GB, 235B is 132GB. The weights should take at most about half the machine's memory, so the context, the prompt cache, and the rest of the machine have the other half. A mixture-of-experts model (the A3B in qwen3.5-35b-a3b) needs memory for all of its parameters and runs at the speed of its active ones.
| Mac | Fits comfortably |
|---|---|
| 24GB | qwen3.5-9b (6GB); gpt-oss-20b (12GB) with little room |
| 36GB | qwen3.5-27b (16GB), gemma-4-26b-a4b (15GB) |
| 64GB | gemma-4-31b (18GB), qwen3.5-35b-a3b (20GB) |
| 256GB | qwen3-235b-a22b-2507 (132GB) |
agency local alias list shows every catalog model's size, agency local list the size of each one you have downloaded, and agency local serve warns when the models it is starting come near the machine's memory.
Prompts cost time up front, replies cost time per token
A reply takes two kinds of time. Reading the prompt is one batch of work that grows with the prompt's length. A 22,000-token prompt took 28 seconds on a 235B model before the first token came out. Writing the reply then costs a fixed time per token. That time is set by how fast the machine can read the model's weights, not by how much arithmetic it does. That is why a 235B model with 22B active parameters and a dense 31B model both write at about 40 tokens per second on the same Mac.
Two things follow. A long prompt is expensive even when the reply is one word, so keep the long prompts to the calls that need them. And the biggest speed-up for a thinking model is not a faster machine, it is thinking: { enabled: false } on the calls that do not need it.
The MLX server also keeps the attention state of its recent prompts and matches a new prompt against them by their longest shared start. A call whose prompt begins the same way as a recent one skips reading that beginning. So put the part that stays the same first, the system prompt and the document, and the part that changes last, the question. Ten questions about one 22,000-token document then pay the 28 seconds once. A conversation gets this for free, because each turn starts with all the turns before it.
Speculative decoding
Since writing a reply is bound by reading the weights, a smaller model can help a larger one. The small model, the draft, guesses the next few tokens cheaply. The large model then checks the whole guess in one pass, which costs it about the same as writing one token, and keeps every token it agrees with. The reply is exactly what the large model would have written alone, only sooner. On prose the gain is usually 1.5x to 2x. On a typed reply it is less, because the grammar makes the draft's guesses wrong more often.
agency local serve mlx:mlx-community/Qwen3-235B-A22B-Instruct-2507-4bit --draft mlx:mlx-community/Qwen3-0.6B-4bit
agency run --local qwen3.5-4b --draft qwen3.5-2b hello.agencyThe first line drafts for an MLX model, the second for a GGUF one. On the server, --draft is written after the model it drafts for, and with several models each can have its own. The draft has to share the main model's tokenizer, which in practice means the smallest member of the same family, and both backends check the pair when the draft loads and refuse one that does not match. --draft-tokens on the server sets how many tokens the draft guesses at a time, four by default. A draft turns off batching on the MLX server, so calls run one at a time there while it is in use. The MLX server also refuses a draft for a model whose attention cache cannot give tokens back, which is every Qwen3.5 and Qwen3-Next model; Qwen3, gpt-oss, and Gemma 4 can take one.
On llama.cpp a drafted model runs greedy unless the call names a temperature. That is the one setting every pair takes: with node-llama-cpp 3.21.1, a Qwen3.5 model that samples on a draft never returns from the call, so the plugin refuses a sampled call on a drafted Qwen3.5 model, and warns once for other families, where a Qwen3 0.6B drafting for the Qwen3 8B returned as usual. llamaCppDraftOptions.allowSampling in the call's metadata turns that check off. The same Qwen3 pair reported no predictions used at any temperature, and ran slower than the 8B alone, so on llama.cpp a draft is not a speed-up today. The draft on the MLX server is the one that can pay.
Whether a draft pays off depends on the pair and the machine, so measure it: run the throughput case of the benchmark with and without the draft and compare the output speed. A draft that is too large gains little, because checking its guesses costs almost what it saves. On the 235B with a 0.6B draft, prose came out slower with the draft than without.
Long prompts and the machine's memory
The MLX server reads a long prompt in chunks. A bigger chunk keeps the GPU busier, so the prompt is read sooner, but the attention scores of one chunk against the whole prompt have to fit in memory at once. Agency sizes the chunk from the machine's memory: 2048 tokens under 64GB, 4096 up to 128GB, and 8192 above. --prefill-step on agency local serve overrides that. Raise it if you have memory to spare and long prompts to read; lower it if the server runs out of memory on a long prompt.
Several calls at once
llama.cpp runs one reply at a time inside your process, and other calls wait their turn. The MLX server batches replies: between one token and the next it looks for new requests, and a request that arrived joins the batch at the next token, so a call never waits for another to finish. Calls that overlap share the GPU and finish sooner together than one after another. Each call still takes about as long as it would alone, so this helps a program with independent calls, not a single slow one. A server with a draft model is the exception; it takes one request at a time. Make the calls concurrent the way you would any Agency work, with fork or parallel.
Time to first token is not measured separately by Agency yet. A hosted model's latency includes the network round trip and the provider's queue, and a local model's does not, so a comparison of latencies alone flatters the local model on short replies.
A type shapes the reply, but cannot make it right
A typed llm() call on a local model gets a reply that fits the type, because a grammar forbids every token that would break it. On a thinking model the grammar waits: the model thinks freely inside its thinking block, and only what comes after it has to fit the type. Two things are worth knowing. A tool call is not JSON, so the grammar cannot be on while the model may still call a tool. On the MLX server, Agency makes a typed call that passes tools as two requests: the model first calls its tools and answers, then a second request asks for that answer in the type, with the grammar on. llama.cpp does the same from smoltalk-llama-cpp 0.8.0, and an older plugin leaves such a call unconstrained. A streamed call is never split, so it is not constrained when it passes tools. And the grammar shapes the reply without judging it. Asked for a date as a string, a model can write March 14, 2026 where you wanted 2026-03-14, and the grammar is satisfied. Say the format you want in the prompt or in the field's description.
The context window is smaller than you think
The prompt, the thinking, and the answer all share one context window. llama.cpp gives a model 32,768 tokens of it in Agency. A prompt near that size leaves no room for a reply, and the call fails with a context error. The MLX server uses the model's own limit, which is usually larger.
llama.cpp and the MLX server, side by side
| llama.cpp | MLX server | |
|---|---|---|
| Runs | inside your process | as a process you start and stop |
| Loads the model | on the first call, every run | once, when you start the server |
| Calls at the same time | one at a time, the rest wait | several, batched together |
| Abandoned reply | stopped at once | stopped at its next token |
| Thinking on, off, budget | yes, where the model's wrapper has a switch | yes |
| Watches for loops | no | yes |
| Speculative decoding | agency run --draft, greedy only | agency local serve --draft |
| Output cap | 16,384 unless the call says | the server's --max-tokens |
| Context window | 32,768 tokens | the model's own |
| Where it runs | any machine | Apple Silicon only |
See also
agency localcommand reference covers every subcommand and config option.- LLM calls covers the
llmfunction itself. - Custom providers covers plugging in providers beyond the built-ins.