local
Use this to manage and run local models. There are two kinds:
- GGUF models run inside the Agency process through llama.cpp. Install
smoltalk-llama-cpponce withnpm i -g smoltalk-llama-cppbefore downloading or running one. - MLX models run in a server on a Mac with Apple Silicon.
agency local servestarts that server for you and puts one port in front of it. You need a Python 3.11 or newer withmlx-lminstalled, andmlx-audioas well for speech models. Agency does not install Python.
Every model has a backend, llama-cpp or mlx. The name you type says which it is:
| You type | Backend |
|---|---|
a curated name such as qwen3.5-2b | from the catalog entry |
hf:org/repo:Q4_K_M or a .gguf path | llama-cpp |
mlx:org/repo or mlx:org/repo@revision | mlx |
a directory holding config.json and .safetensors files | mlx |
agency local download # pick a model from the catalog interactively
agency local download qwen3.5-2b # curated name, alias, hf: URI, or .gguf path
agency local list # the catalog, with each model's kind and backend, downloaded models marked
agency local list -l # ...and each model's description
agency local list --kind image # one kind: chat, embedding, speech, image, vision, or controlnet
agency local remove my7b # remove the alias, keep the files
agency local remove my7b -f # remove the alias and delete the files
agency local resolve my7b # show the backend and what a name/alias maps to
agency local alias add my7b hf:Qwen/Qwen2.5-7B-Instruct-GGUF:Q4_K_M
agency local alias add coder /models/mlx-community--Qwen3-Coder-Next-4bit # an MLX model directory
agency local alias list # curated + your aliases, with descriptions
agency local alias remove my7bThe agent and agency run have shortcuts for the common case:
agency agent --local qwen3.5-2b # download (if needed) + run the agent locally
agency run --local qwen3.5-2b my.agency # download (if needed) + run a program locally
agency run --local coder my.agency # an MLX model: needs `agency local serve coder` runningThe agent's --local runs the local model as both the fast and slow model, so the deep subagents stay local too; it ignores --model/--fastmodel/--slowmodel. On agency run, --local and --model are mutually exclusive. See the local models guide for a walkthrough.
Running an MLX model
An MLX model is a Hugging Face repo of .safetensors files, such as mlx-community/Qwen3-Coder-Next-4bit. Agency runs it in a separate process, which agency local serve starts:
# Terminal 1.
agency local download mlx:mlx-community/Qwen3-Coder-Next-4bit
agency local serve mlx:mlx-community/Qwen3-Coder-Next-4bit
# Terminal 2.
agency run --local mlx:mlx-community/Qwen3-Coder-Next-4bit my.agencyServing needs a Python with mlx-lm installed. Agency looks for it at --python, then client.mlx.python in agency.json, then AGENCY_MLX_PYTHON, then ~/.agency-agent/mlx-env/bin/python. When it cannot find one, it prints the commands to create that environment:
python3.12 -m venv ~/.agency-agent/mlx-env
~/.agency-agent/mlx-env/bin/pip install mlx-lm==0.31.3 llguidance==1.8.0Serve and run must name the model the same way, which is the string agency local resolve prints: the repo id for an mlx: URI, or the directory path for a directory alias. A request naming a model the server was not started with gets a 404 that says how to start it.
The server listens on http://127.0.0.1:8080/v1 by default. For another port, pass --port and set MLX_BASE_URL or client.baseUrl.mlx in agency.json to match.
A Hugging Face cache directory works as is, and you can name either the repo folder or one snapshot inside it:
agency local serve /Volumes/models/hf/hub/models--mlx-community--Qwen3.8-27B-4bit
agency local serve /Volumes/models/hf/hub/models--mlx-community--Qwen3.8-27B-4bit/snapshots/3e6447f0The repo folder resolves through its refs/main to the snapshot that ref names; with no ref and one snapshot, that snapshot is used; with several and no ref, Agency lists them and asks you to name one. A snapshot's entries are symlinks into blobs/, and Agency follows them for this one check.
Use --model-dir to point the whole local command at another directory for one run — your own models directory, or somebody else's Hugging Face cache:
agency local --model-dir /Volumes/models/hf/hub list
agency local --model-dir /Volumes/models/hf/hub serve # a picker over everything thereIt takes precedence over AGENCY_MODELS_DIR and client.modelsDir, and applies to list, download, serve and remove alike. Agency reads a Hugging Face cache but never writes one: agency local download --model-dir <hub> puts the model in Agency's own layout beside it, and agency local remove -f refuses a cache model, since deleting a snapshot of symlinks would leave the bytes in blobs/ behind.
Naming a model
The mlx: prefix names a model by its Hugging Face repo id, wherever its files are — Agency's models directory, or a Hugging Face cache. agency local serve mlx:mlx-community/Qwen3.8-27B-4bit serves it under that repo id, whichever layout holds it, and agency run --local mlx:mlx-community/Qwen3.8-27B-4bit sends the same name. If you already have the model, the prefix is optional: a bare mlx-community/Qwen3.8-27B-4bit works too. It is required only for a model you have not downloaded yet, which is why messages print that spelling.
To download one, use its mlx: URI. The download runs as parallel byte-range requests, resumes if interrupted, and verifies each file: against the SHA-256 Hugging Face publishes for large (LFS) files, and by size for small plain files such as config.json:
agency local download mlx:mlx-community/Qwen3-Coder-Next-4bit
agency local download mlx:mlx-community/Qwen3-Coder-Next-4bit@7b9321e # a pinned commitServing models
agency local serve starts the servers for you and puts one port in front of them:
agency local serve # pick from what you have downloaded
agency local serve mlx:mlx-community/Qwen3-Coder-Next-4bit coder
agency local serve --port 8080 --verbose coder # ...and log every request and reply in fullModels that are not chat models need a flag, because they answer a different route and run a different program:
agency local serve coder --embedding qwen3-embedding-4b-mlx # /v1/embeddings
agency local serve --speech qwen3-tts-mlx # /v1/audio/speech
agency local serve --speech orpheus-3b-mlx --speech qwen3-tts-mlxBoth flags are repeatable, and either can be combined with chat models in the same command. A speech model needs mlx-audio in the same Python:
~/.agency-agent/mlx-env/bin/pip install mlx-audio==0.5.4Speech answers POST /v1/audio/speech in the OpenAI shape, returning WAV or raw PCM:
curl -s http://127.0.0.1:8080/v1/audio/speech -H 'content-type: application/json' \
-d '{"model": "mlx-community/Qwen3-TTS-12Hz-1.7B-CustomVoice-8bit", "input": "Hello there.", "instructions": "Calm."}' \
-o hello.wavThe catalog knows which models are which, so agency local serve qwen3-tts-mlx starts the speech server with no flag, and --embedding qwen3-tts-mlx is refused before anything loads, naming the command that works. Some models load a second repo by name: agency local download orpheus-3b-mlx fetches its audio decoder as well, which is what lets it start with no network.
A model that does not fit in memory beside the others can be served on demand with --lazy:
agency local serve florence-2 --lazy z-image-turbo --lazy flux2-klein-9bThis serves three models. florence-2 loads before the port opens and stays loaded, as every model does without the flag. The two image models load on the first request that names each of them, and that request waits for the load. When a lazy model does not fit, the server stops the lazy model that has been idle longest to make room. A model that is not lazy is never stopped for one that is. All three names appear in GET /v1/models as soon as the port opens, and GET /v1/agency/status says which are loaded.
--lazy names a model, so a lazy model is not also written as a plain argument; z-image-turbo --lazy z-image-turbo is refused. To force a kind, give the kind flag as well: --vlm qwen3.5-9b --lazy qwen3.5-9b. A --draft written after --lazy <model> belongs to that model.
The decision whether a model fits is an estimate: its size on disk, plus 4 GB for an image model or 1 GB for any other kind, against the memory available now less a small reserve. AGENCY_ALLOW_MEMORY_OVERCOMMIT=1 loads a lazy model even when the estimate says no, for a machine where it is too cautious.
Three routes under /v1/agency/ are about the server itself. GET /v1/agency/status lists each model's state. POST /v1/agency/cancel with {"model": "..."} ends every request running on that model. POST /v1/agency/shutdown stops the server, which then exits as it does on Ctrl-C. The two POST routes take content-type: application/json, and all three answer requests made to 127.0.0.1 or localhost only.
Image models run on diffusers rather than MLX. Any diffusers model is served as an image model, with or without --image:
agency local download z-image-turbo
agency local serve --image z-image-turbo # /v1/images/generationsAn image model is named by a catalog name, a diffusers: URI such as diffusers:Tongyi-MAI/Z-Image-Turbo, or a directory holding model_index.json. It needs torch and diffusers in the same Python, and no MLX packages:
~/.agency-agent/mlx-env/bin/pip install torch==2.14.0 diffusers==0.40.0 transformers==5.17.0 accelerate==1.15.0 sentencepiece==0.2.2 protobuf==7.36.2Agency code calls it with generateImageLocal from std::image. On an M5 Ultra, z-image-turbo makes a 1024×1024 image in about 8 seconds and chroma1-hd in about 90. The catalog also has qwen-image-2512, for legible text inside an image, and flux2-klein-4b, a small model for Macs with less memory; neither has been timed yet.
An SDXL model, such as the illustration finetune NoobAI-XL, is named by its repo, and takes LoRA adapters you trained or downloaded. Put the .safetensors files in a folder and name the folder in agency.json:
agency local download diffusers:Laxhar/noobai-XL-1.1
agency local serve diffusers:Laxhar/noobai-XL-1.1{ "client": { "adaptersDir": "./adapters" } }A request then asks for an adapter by its file name without the extension: generateImageLocal("sketch, a cat on a chair", "diffusers:Laxhar/noobai-XL-1.1", lora: "sketch") loads ./adapters/sketch.safetensors the first time and keeps it. Only a request that names an adapter gets one. A file dropped into the folder is usable with no restart, and a file trained again is read again the next time a request names it. A relative adaptersDir is taken from the folder agency.json is in.
A ControlNet is downloaded into client.controlnetsDir rather than the models directory, since an image server loads it from there when a call names it: agency local download controlnet-scribble-sdxl, then generateImageLocal(..., controlnet: "controlnet-scribble-sdxl", controlImage: "./pose.png"). It is never served on its own.
Vision models find things in images rather than making them. Two are in the catalog: wd14-tagger describes a picture as booru tags, and florence-2 finds the objects you name, tags, and captions. They are served the same way and called from std::vision:
agency local download wd14-tagger florence-2
agency local serve wd14-tagger florence-2import { detectObjects, tagImage } from "std::vision"
const found = detectObjects("page.png", ["person", "desk"], "florence-2") with approve
const tags = tagImage("page.png", "wd14-tagger") with approveThe image is read on this machine and never sent anywhere; each call raises one std::vision effect naming the file, the task, and the model.
Two more vision models find a particular thing rather than any thing of its kind. owlv2-base finds the objects you name in one pass, and findRegions asks it for a box around every thing in a picture. dinov2-base answers embedImage, which turns a picture, or boxes in it, into embeddings you can compare. The guide's "Find your own things" section shows the two together.
It prints a line for every request that reaches it:
POST /v1/chat/completions mlx-community/Qwen3.5-4B-MLX-4bit 200 2.6s 16→129 tokThat is the method and path, the model the request named, the status, how long it took, and the prompt and completion tokens the server reported. With --verbose (or --log-prompts, the same thing), the whole request body (→) and the reply (←) follow underneath: indented JSON for a normal reply, and the data: frames as they came for a streamed one. The server keeps the first 1 MiB of a reply for the log and marks a longer one … (truncated). What the client receives is never changed. Color is on when the output is a terminal; NO_COLOR turns it off and FORCE_COLOR turns it on for a log file.
A gated repo needs HF_TOKEN set to a token that has accepted its terms. The token is sent to huggingface.co only and never stored. client.mlx.downloadConcurrency in agency.json sets the number of parallel requests (default 8).
Subcommands
| command | purpose |
|---|---|
agency local --model-dir <path> <subcommand> | Use <path> as the models directory for this run, ahead of AGENCY_MODELS_DIR and client.modelsDir. A Hugging Face cache under it is read too, so pointing at hf/hub lists and serves what another tool downloaded. |
agency local list [--kind <kind>] | Show the full catalog with each model's kind and backend, and a checkmark and on-disk size for downloaded models, in kind order: chat, embedding, speech, image. --kind shows one kind. The first line names the models directory. Files that match no catalog entry appear under OTHER FILES with their kind. Add -l / --long to print each model's description on its own line below its row. Works without smoltalk-llama-cpp installed. |
agency local download [value] [--kind <kind>] | Download a model if not already cached; prints the source it resolved to and the local path. The download records what kind of model it is (chat, embedding, speech, or image) from the catalog or from its files, so serve needs no flag; --kind says it for a repo whose files do not. --kind is refused for a directory or a .gguf file, which keep no record; name the kind when serving a directory instead. <value> may be a curated short name, an alias, an hf: URI, or an existing .gguf path. With no value, opens an interactive picker (in scripts it prints the catalog and exits 1 instead). An mlx: URI downloads the whole repo into <modelsDir>/mlx/<org>--<repo>, resuming if interrupted. A diffusers: URI downloads into <modelsDir>/diffusers/<org>--<repo>, keeping only the files the image pipeline reads. A model directory is returned as is. |
agency local serve [model [--draft <model>] [--draft-tokens 4]]... [--embedding <model>] [--speech <model>] [--image <model>] [--port 8080] [--max-tokens 16384] [--reasoning-budget <n>] [--hedge-limit 12] [--repeat-limit 3] [--limit-answers] [--prefill-step <n>] [--python <path>] [--verbose] | Serve one or more MLX or diffusers models in this terminal. Starts one process per model, waits until each has loaded, then listens on --port. Any model is named on its own: serve reads its kind (chat, embedding, speech, image, or vision) from the catalog, the download's record, or its files, and starts the right server for it, so an embedding model named alone serves /v1/embeddings, a speech model /v1/audio/speech, and an image model /v1/images/generations. --embedding, --speech, and --image are repeatable and give a model's kind yourself. Use one when the files say nothing, or say the wrong thing: a flag outranks the record and the files. A flag the catalog disagrees with is refused before anything loads, with the command that works. A model whose kind nothing gives, named without a flag, is refused with the flags to use. --draft and --draft-tokens are per-model options: written after a chat model, they apply to that model alone, so serve big --draft small other drafts for big and not other. A --draft before every model, or after a model that is not a chat model, is refused, and so is --draft-tokens with no --draft for its model; serve --draft <model> with no model to serve no longer opens the picker. With no model, opens a picker over the models you have downloaded whose kind is known (in scripts it lists them and exits 1 instead). A request for a model you did not name gets a 404 naming the command to start it. It never downloads. An mlx: model must be downloaded first, and a directory works as is. Every request is logged as one line, with audio and image replies logged by size; --verbose (or --log-prompts) adds the whole request and reply bodies. --max-tokens is a ceiling: a request asking for a longer reply gets one no longer than that. A chat server may spend a sixteenth of the machine's memory holding the attention state of its recent prompts, and a reply whose client has gone is stopped at its next token. A reply that goes in circles is cut short: thinking past a budget of half max_tokens, or, in its thinking, twelve "But wait" style second thoughts or a sentence repeated three times within the last two thousand tokens; --limit-answers watches answers the same way (see the guide on local models). Ctrl-C stops everything. |
agency local remove <name> [-f] | Remove the alias for a model and keep its files, printing where they are. With -f, delete the files too: the .gguf file, or the whole MLX or diffusers model directory. Files outside the models directory are never deleted. |
agency local resolve <value> | Show the backend and what a name/alias maps to, without downloading. |
agency local refresh [url] | Fetch the remote model catalog and update the source:"remote" aliases in agency.json. Adds/updates models from the catalog, removes ones it dropped, and skips any name you've aliased yourself (printing what it would have set). |
agency local alias list | List usable short names. Curated entries show params, tags, size, context window, and license (with the description on the next line); your aliases show their target. |
agency local alias add <name> <target> | Add a short-name alias. The target is an hf: URI, a .gguf path, an mlx: or diffusers: URI, or a model directory. Prints the agency.json path that was edited. |
agency local alias remove <name> | Remove a short-name alias. Prints the agency.json path that was inspected (the file is left untouched if the alias wasn't present). |
Refreshing the catalog
agency local refresh pulls a JSON catalog of recommended models and writes them into client.modelAliases as rich, source:"remote"-tagged entries, so new model recommendations arrive without upgrading agency. Your own hand-added aliases are never overwritten — on a name clash the command keeps yours and prints the remote value it skipped.
URL resolution (first wins): the [url] argument, AGENCY_MODEL_CATALOG_URL, client.modelCatalogUrl in agency.json, then the built-in default (raw.githubusercontent.com/egonSchiele/agency-lang/main/packages/agency-lang/data/model-catalog.json). A remote URL must be https:// (an http:// source is rejected). The source may also be a local file path or file:// URL — e.g. agency local refresh ./my-catalog.json — which reads the catalog from disk without any network call.
Heads-up: the first refresh writes one tagged entry per catalog model into
client.modelAliases, so a freshly-refreshedagency.jsonwill be noticeably larger than before. The entries are tagged with"source": "remote"— anything without that tag (your own aliases) is never touched. Re-runningagency local refreshoverwrites only thesource:"remote"entries.
Where things live
- Cache dir:
AGENCY_MODELS_DIRenv var, elseclient.modelsDirin the nearestagency.json, else~/.agency-agent/models.agency local listprints the resolved directory on its first line. The default is shared withagency agent --localandagency run --local, so alocal downloadpre-populates what both reuse. - Aliases: written to the nearest
agency.jsonwalking up from the current directory; if none is found,~/agency.jsonis used. The CLI prints which file it edited on every add/remove. - Curated catalog: permissive licenses only (apache-2.0 / mit); restrictively-licensed weights (Gemma 1–3's custom terms, llama) are intentionally excluded. Gemma 4 ships under apache-2.0, so it is included.
Config
Aliases, the models cache dir, and the catalog URL live under client and are read at runtime, so edits take effect on the next call:
{
"client": {
"modelAliases": {
"my7b": "hf:Qwen/Qwen2.5-7B-Instruct-GGUF:Q4_K_M",
"coder": "mlx:mlx-community/Qwen3-Coder-Next-4bit",
// The object form needs a "backend" that matches the uri.
"big": { "backend": "mlx", "uri": "mlx:mlx-community/DeepSeek-V4-Flash-4bit", "params": "300B" },
// "kind" says what the model takes and returns, and "tags" says what it is good for.
"embedder": { "backend": "mlx", "uri": "mlx:mlx-community/Qwen3-Embedding-4B-4bit-DWQ", "kind": "embedding" },
"thinker": { "backend": "mlx", "uri": "mlx:org/some-model", "kind": "chat", "tags": ["reasoning", "coding"] }
},
"modelsDir": "/data/agency-models",
// Override the URL `agency local refresh` fetches the model catalog from.
// Defaults to the catalog committed in the agency repo.
"modelCatalogUrl": "https://example.com/my-model-catalog.json"
}
}An object alias can also say what kind of model it is. kind is one of chat, embedding, speech, or image, and agency local serve uses it to pick the server, so an alias for an embedding model should say "kind": "embedding". tags is a list of words for what the model is good at, such as ["reasoning", "coding"], and shows in the TAGS column of agency local list. Both are optional. A kind that is not one of the four, or tags that are not a list of strings, is an error when the config is read. A catalog you host through modelCatalogUrl uses the same two fields.
See also
- Using local models guide — the walkthrough, from install to
llm()calls in code. agency agent --local— the easy button that composes the local-model primitives.- Custom providers guide — for using any other provider.
Chat with images
agency local serve --vlm qwen3.5-9b-mlx starts the model with mlx-vlm==0.7.0. The flag is repeatable and required for image-input chat. Without it, the model uses the text runtime. See chat with pictures for installation and an Agency example.
A shared serve option such as --hedge-limit is refused when no model in the command supports it. It is allowed when a text model in the same command uses it. --draft applies to the preceding model and is refused for a model named with --vlm.