Skip to content

local ​

Use this to manage and run local models. There are two kinds:

  • GGUF models run inside the Agency process through llama.cpp. Install smoltalk-llama-cpp once with npm i -g smoltalk-llama-cpp before downloading or running one.
  • MLX models run in a server on a Mac with Apple Silicon. agency local serve starts that server for you and puts one port in front of it. You need a Python 3.11 or newer with mlx-lm installed, and mlx-audio as well for speech models. Agency does not install Python.

Every model has a backend, llama-cpp or mlx. The name you type says which it is:

You typeBackend
a curated name such as qwen3.5-2bfrom the catalog entry
hf:org/repo:Q4_K_M or a .gguf pathllama-cpp
mlx:org/repo or mlx:org/repo@revisionmlx
a directory holding config.json and .safetensors filesmlx
bash
agency local download                     # pick a model from the catalog interactively
agency local download qwen3.5-2b          # curated name, alias, hf: URI, or .gguf path
agency local list                         # the catalog, with each model's kind and backend, downloaded models marked
agency local list -l                      # ...and each model's description
agency local list --kind image            # one kind: chat, embedding, speech, image, vision, or controlnet
agency local remove my7b                  # remove the alias, keep the files
agency local remove my7b -f               # remove the alias and delete the files
agency local resolve my7b                 # show the backend and what a name/alias maps to

agency local alias add my7b hf:Qwen/Qwen2.5-7B-Instruct-GGUF:Q4_K_M
agency local alias add coder /models/mlx-community--Qwen3-Coder-Next-4bit   # an MLX model directory
agency local alias list                   # curated + your aliases, with descriptions
agency local alias remove my7b

The agent and agency run have shortcuts for the common case:

bash
agency agent --local qwen3.5-2b           # download (if needed) + run the agent locally
agency run --local qwen3.5-2b my.agency   # download (if needed) + run a program locally
agency run --local coder my.agency        # an MLX model: needs `agency local serve coder` running

The agent's --local runs the local model as both the fast and slow model, so the deep subagents stay local too; it ignores --model/--fastmodel/--slowmodel. On agency run, --local and --model are mutually exclusive. See the local models guide for a walkthrough.

Running an MLX model ​

An MLX model is a Hugging Face repo of .safetensors files, such as mlx-community/Qwen3-Coder-Next-4bit. Agency runs it in a separate process, which agency local serve starts:

bash
# Terminal 1.
agency local download mlx:mlx-community/Qwen3-Coder-Next-4bit
agency local serve mlx:mlx-community/Qwen3-Coder-Next-4bit

# Terminal 2.
agency run --local mlx:mlx-community/Qwen3-Coder-Next-4bit my.agency

Serving needs a Python with mlx-lm installed. Agency looks for it at --python, then client.mlx.python in agency.json, then AGENCY_MLX_PYTHON, then ~/.agency-agent/mlx-env/bin/python. When it cannot find one, it prints the commands to create that environment:

bash
python3.12 -m venv ~/.agency-agent/mlx-env
~/.agency-agent/mlx-env/bin/pip install mlx-lm==0.31.3 llguidance==1.8.0

Serve and run must name the model the same way, which is the string agency local resolve prints: the repo id for an mlx: URI, or the directory path for a directory alias. A request naming a model the server was not started with gets a 404 that says how to start it.

The server listens on http://127.0.0.1:8080/v1 by default. For another port, pass --port and set MLX_BASE_URL or client.baseUrl.mlx in agency.json to match.

A Hugging Face cache directory works as is, and you can name either the repo folder or one snapshot inside it:

bash
agency local serve /Volumes/models/hf/hub/models--mlx-community--Qwen3.8-27B-4bit
agency local serve /Volumes/models/hf/hub/models--mlx-community--Qwen3.8-27B-4bit/snapshots/3e6447f0

The repo folder resolves through its refs/main to the snapshot that ref names; with no ref and one snapshot, that snapshot is used; with several and no ref, Agency lists them and asks you to name one. A snapshot's entries are symlinks into blobs/, and Agency follows them for this one check.

Use --model-dir to point the whole local command at another directory for one run — your own models directory, or somebody else's Hugging Face cache:

bash
agency local --model-dir /Volumes/models/hf/hub list
agency local --model-dir /Volumes/models/hf/hub serve      # a picker over everything there

It takes precedence over AGENCY_MODELS_DIR and client.modelsDir, and applies to list, download, serve and remove alike. Agency reads a Hugging Face cache but never writes one: agency local download --model-dir <hub> puts the model in Agency's own layout beside it, and agency local remove -f refuses a cache model, since deleting a snapshot of symlinks would leave the bytes in blobs/ behind.

Naming a model ​

The mlx: prefix names a model by its Hugging Face repo id, wherever its files are — Agency's models directory, or a Hugging Face cache. agency local serve mlx:mlx-community/Qwen3.8-27B-4bit serves it under that repo id, whichever layout holds it, and agency run --local mlx:mlx-community/Qwen3.8-27B-4bit sends the same name. If you already have the model, the prefix is optional: a bare mlx-community/Qwen3.8-27B-4bit works too. It is required only for a model you have not downloaded yet, which is why messages print that spelling.

To download one, use its mlx: URI. The download runs as parallel byte-range requests, resumes if interrupted, and verifies each file: against the SHA-256 Hugging Face publishes for large (LFS) files, and by size for small plain files such as config.json:

bash
agency local download mlx:mlx-community/Qwen3-Coder-Next-4bit
agency local download mlx:mlx-community/Qwen3-Coder-Next-4bit@7b9321e   # a pinned commit

Serving models ​

agency local serve starts the servers for you and puts one port in front of them:

bash
agency local serve                                     # pick from what you have downloaded
agency local serve mlx:mlx-community/Qwen3-Coder-Next-4bit coder
agency local serve --port 8080 --verbose coder         # ...and log every request and reply in full

Models that are not chat models need a flag, because they answer a different route and run a different program:

bash
agency local serve coder --embedding qwen3-embedding-4b-mlx   # /v1/embeddings
agency local serve --speech qwen3-tts-mlx                     # /v1/audio/speech
agency local serve --speech orpheus-3b-mlx --speech qwen3-tts-mlx

Both flags are repeatable, and either can be combined with chat models in the same command. A speech model needs mlx-audio in the same Python:

bash
~/.agency-agent/mlx-env/bin/pip install mlx-audio==0.5.4

Speech answers POST /v1/audio/speech in the OpenAI shape, returning WAV or raw PCM:

bash
curl -s http://127.0.0.1:8080/v1/audio/speech -H 'content-type: application/json' \
  -d '{"model": "mlx-community/Qwen3-TTS-12Hz-1.7B-CustomVoice-8bit", "input": "Hello there.", "instructions": "Calm."}' \
  -o hello.wav

The catalog knows which models are which, so agency local serve qwen3-tts-mlx starts the speech server with no flag, and --embedding qwen3-tts-mlx is refused before anything loads, naming the command that works. Some models load a second repo by name: agency local download orpheus-3b-mlx fetches its audio decoder as well, which is what lets it start with no network.

A model that does not fit in memory beside the others can be served on demand with --lazy:

bash
agency local serve florence-2 --lazy z-image-turbo --lazy flux2-klein-9b

This serves three models. florence-2 loads before the port opens and stays loaded, as every model does without the flag. The two image models load on the first request that names each of them, and that request waits for the load. When a lazy model does not fit, the server stops the lazy model that has been idle longest to make room. A model that is not lazy is never stopped for one that is. All three names appear in GET /v1/models as soon as the port opens, and GET /v1/agency/status says which are loaded.

--lazy names a model, so a lazy model is not also written as a plain argument; z-image-turbo --lazy z-image-turbo is refused. To force a kind, give the kind flag as well: --vlm qwen3.5-9b --lazy qwen3.5-9b. A --draft written after --lazy <model> belongs to that model.

The decision whether a model fits is an estimate: its size on disk, plus 4 GB for an image model or 1 GB for any other kind, against the memory available now less a small reserve. AGENCY_ALLOW_MEMORY_OVERCOMMIT=1 loads a lazy model even when the estimate says no, for a machine where it is too cautious.

Three routes under /v1/agency/ are about the server itself. GET /v1/agency/status lists each model's state. POST /v1/agency/cancel with {"model": "..."} ends every request running on that model. POST /v1/agency/shutdown stops the server, which then exits as it does on Ctrl-C. The two POST routes take content-type: application/json, and all three answer requests made to 127.0.0.1 or localhost only.

Image models run on diffusers rather than MLX. Any diffusers model is served as an image model, with or without --image:

bash
agency local download z-image-turbo
agency local serve --image z-image-turbo                      # /v1/images/generations

An image model is named by a catalog name, a diffusers: URI such as diffusers:Tongyi-MAI/Z-Image-Turbo, or a directory holding model_index.json. It needs torch and diffusers in the same Python, and no MLX packages:

bash
~/.agency-agent/mlx-env/bin/pip install torch==2.14.0 diffusers==0.40.0 transformers==5.17.0 accelerate==1.15.0 sentencepiece==0.2.2 protobuf==7.36.2

Agency code calls it with generateImageLocal from std::image. On an M5 Ultra, z-image-turbo makes a 1024×1024 image in about 8 seconds and chroma1-hd in about 90. The catalog also has qwen-image-2512, for legible text inside an image, and flux2-klein-4b, a small model for Macs with less memory; neither has been timed yet.

An SDXL model, such as the illustration finetune NoobAI-XL, is named by its repo, and takes LoRA adapters you trained or downloaded. Put the .safetensors files in a folder and name the folder in agency.json:

bash
agency local download diffusers:Laxhar/noobai-XL-1.1
agency local serve diffusers:Laxhar/noobai-XL-1.1
json
{ "client": { "adaptersDir": "./adapters" } }

A request then asks for an adapter by its file name without the extension: generateImageLocal("sketch, a cat on a chair", "diffusers:Laxhar/noobai-XL-1.1", lora: "sketch") loads ./adapters/sketch.safetensors the first time and keeps it. Only a request that names an adapter gets one. A file dropped into the folder is usable with no restart, and a file trained again is read again the next time a request names it. A relative adaptersDir is taken from the folder agency.json is in.

A ControlNet is downloaded into client.controlnetsDir rather than the models directory, since an image server loads it from there when a call names it: agency local download controlnet-scribble-sdxl, then generateImageLocal(..., controlnet: "controlnet-scribble-sdxl", controlImage: "./pose.png"). It is never served on its own.

Vision models find things in images rather than making them. Two are in the catalog: wd14-tagger describes a picture as booru tags, and florence-2 finds the objects you name, tags, and captions. They are served the same way and called from std::vision:

bash
agency local download wd14-tagger florence-2
agency local serve wd14-tagger florence-2
ts
import { detectObjects, tagImage } from "std::vision"
const found = detectObjects("page.png", ["person", "desk"], "florence-2") with approve
const tags = tagImage("page.png", "wd14-tagger") with approve

The image is read on this machine and never sent anywhere; each call raises one std::vision effect naming the file, the task, and the model.

Two more vision models find a particular thing rather than any thing of its kind. owlv2-base finds the objects you name in one pass, and findRegions asks it for a box around every thing in a picture. dinov2-base answers embedImage, which turns a picture, or boxes in it, into embeddings you can compare. The guide's "Find your own things" section shows the two together.

It prints a line for every request that reaches it:

POST /v1/chat/completions  mlx-community/Qwen3.5-4B-MLX-4bit  200  2.6s  16→129 tok

That is the method and path, the model the request named, the status, how long it took, and the prompt and completion tokens the server reported. With --verbose (or --log-prompts, the same thing), the whole request body (→) and the reply (←) follow underneath: indented JSON for a normal reply, and the data: frames as they came for a streamed one. The server keeps the first 1 MiB of a reply for the log and marks a longer one … (truncated). What the client receives is never changed. Color is on when the output is a terminal; NO_COLOR turns it off and FORCE_COLOR turns it on for a log file.

A gated repo needs HF_TOKEN set to a token that has accepted its terms. The token is sent to huggingface.co only and never stored. client.mlx.downloadConcurrency in agency.json sets the number of parallel requests (default 8).

Subcommands ​

commandpurpose
agency local --model-dir <path> <subcommand>Use <path> as the models directory for this run, ahead of AGENCY_MODELS_DIR and client.modelsDir. A Hugging Face cache under it is read too, so pointing at hf/hub lists and serves what another tool downloaded.
agency local list [--kind <kind>]Show the full catalog with each model's kind and backend, and a checkmark and on-disk size for downloaded models, in kind order: chat, embedding, speech, image. --kind shows one kind. The first line names the models directory. Files that match no catalog entry appear under OTHER FILES with their kind. Add -l / --long to print each model's description on its own line below its row. Works without smoltalk-llama-cpp installed.
agency local download [value] [--kind <kind>]Download a model if not already cached; prints the source it resolved to and the local path. The download records what kind of model it is (chat, embedding, speech, or image) from the catalog or from its files, so serve needs no flag; --kind says it for a repo whose files do not. --kind is refused for a directory or a .gguf file, which keep no record; name the kind when serving a directory instead. <value> may be a curated short name, an alias, an hf: URI, or an existing .gguf path. With no value, opens an interactive picker (in scripts it prints the catalog and exits 1 instead). An mlx: URI downloads the whole repo into <modelsDir>/mlx/<org>--<repo>, resuming if interrupted. A diffusers: URI downloads into <modelsDir>/diffusers/<org>--<repo>, keeping only the files the image pipeline reads. A model directory is returned as is.
agency local serve [model [--draft <model>] [--draft-tokens 4]]... [--embedding <model>] [--speech <model>] [--image <model>] [--port 8080] [--max-tokens 16384] [--reasoning-budget <n>] [--hedge-limit 12] [--repeat-limit 3] [--limit-answers] [--prefill-step <n>] [--python <path>] [--verbose]Serve one or more MLX or diffusers models in this terminal. Starts one process per model, waits until each has loaded, then listens on --port. Any model is named on its own: serve reads its kind (chat, embedding, speech, image, or vision) from the catalog, the download's record, or its files, and starts the right server for it, so an embedding model named alone serves /v1/embeddings, a speech model /v1/audio/speech, and an image model /v1/images/generations. --embedding, --speech, and --image are repeatable and give a model's kind yourself. Use one when the files say nothing, or say the wrong thing: a flag outranks the record and the files. A flag the catalog disagrees with is refused before anything loads, with the command that works. A model whose kind nothing gives, named without a flag, is refused with the flags to use. --draft and --draft-tokens are per-model options: written after a chat model, they apply to that model alone, so serve big --draft small other drafts for big and not other. A --draft before every model, or after a model that is not a chat model, is refused, and so is --draft-tokens with no --draft for its model; serve --draft <model> with no model to serve no longer opens the picker. With no model, opens a picker over the models you have downloaded whose kind is known (in scripts it lists them and exits 1 instead). A request for a model you did not name gets a 404 naming the command to start it. It never downloads. An mlx: model must be downloaded first, and a directory works as is. Every request is logged as one line, with audio and image replies logged by size; --verbose (or --log-prompts) adds the whole request and reply bodies. --max-tokens is a ceiling: a request asking for a longer reply gets one no longer than that. A chat server may spend a sixteenth of the machine's memory holding the attention state of its recent prompts, and a reply whose client has gone is stopped at its next token. A reply that goes in circles is cut short: thinking past a budget of half max_tokens, or, in its thinking, twelve "But wait" style second thoughts or a sentence repeated three times within the last two thousand tokens; --limit-answers watches answers the same way (see the guide on local models). Ctrl-C stops everything.
agency local remove <name> [-f]Remove the alias for a model and keep its files, printing where they are. With -f, delete the files too: the .gguf file, or the whole MLX or diffusers model directory. Files outside the models directory are never deleted.
agency local resolve <value>Show the backend and what a name/alias maps to, without downloading.
agency local refresh [url]Fetch the remote model catalog and update the source:"remote" aliases in agency.json. Adds/updates models from the catalog, removes ones it dropped, and skips any name you've aliased yourself (printing what it would have set).
agency local alias listList usable short names. Curated entries show params, tags, size, context window, and license (with the description on the next line); your aliases show their target.
agency local alias add <name> <target>Add a short-name alias. The target is an hf: URI, a .gguf path, an mlx: or diffusers: URI, or a model directory. Prints the agency.json path that was edited.
agency local alias remove <name>Remove a short-name alias. Prints the agency.json path that was inspected (the file is left untouched if the alias wasn't present).

Refreshing the catalog ​

agency local refresh pulls a JSON catalog of recommended models and writes them into client.modelAliases as rich, source:"remote"-tagged entries, so new model recommendations arrive without upgrading agency. Your own hand-added aliases are never overwritten — on a name clash the command keeps yours and prints the remote value it skipped.

URL resolution (first wins): the [url] argument, AGENCY_MODEL_CATALOG_URL, client.modelCatalogUrl in agency.json, then the built-in default (raw.githubusercontent.com/egonSchiele/agency-lang/main/packages/agency-lang/data/model-catalog.json). A remote URL must be https:// (an http:// source is rejected). The source may also be a local file path or file:// URL — e.g. agency local refresh ./my-catalog.json — which reads the catalog from disk without any network call.

Heads-up: the first refresh writes one tagged entry per catalog model into client.modelAliases, so a freshly-refreshed agency.json will be noticeably larger than before. The entries are tagged with "source": "remote" — anything without that tag (your own aliases) is never touched. Re-running agency local refresh overwrites only the source:"remote" entries.

Where things live ​

  • Cache dir: AGENCY_MODELS_DIR env var, else client.modelsDir in the nearest agency.json, else ~/.agency-agent/models. agency local list prints the resolved directory on its first line. The default is shared with agency agent --local and agency run --local, so a local download pre-populates what both reuse.
  • Aliases: written to the nearest agency.json walking up from the current directory; if none is found, ~/agency.json is used. The CLI prints which file it edited on every add/remove.
  • Curated catalog: permissive licenses only (apache-2.0 / mit); restrictively-licensed weights (Gemma 1–3's custom terms, llama) are intentionally excluded. Gemma 4 ships under apache-2.0, so it is included.

Config ​

Aliases, the models cache dir, and the catalog URL live under client and are read at runtime, so edits take effect on the next call:

jsonc
{
  "client": {
    "modelAliases": {
      "my7b": "hf:Qwen/Qwen2.5-7B-Instruct-GGUF:Q4_K_M",
      "coder": "mlx:mlx-community/Qwen3-Coder-Next-4bit",
      // The object form needs a "backend" that matches the uri.
      "big": { "backend": "mlx", "uri": "mlx:mlx-community/DeepSeek-V4-Flash-4bit", "params": "300B" },
      // "kind" says what the model takes and returns, and "tags" says what it is good for.
      "embedder": { "backend": "mlx", "uri": "mlx:mlx-community/Qwen3-Embedding-4B-4bit-DWQ", "kind": "embedding" },
      "thinker": { "backend": "mlx", "uri": "mlx:org/some-model", "kind": "chat", "tags": ["reasoning", "coding"] }
    },
    "modelsDir": "/data/agency-models",
    // Override the URL `agency local refresh` fetches the model catalog from.
    // Defaults to the catalog committed in the agency repo.
    "modelCatalogUrl": "https://example.com/my-model-catalog.json"
  }
}

An object alias can also say what kind of model it is. kind is one of chat, embedding, speech, or image, and agency local serve uses it to pick the server, so an alias for an embedding model should say "kind": "embedding". tags is a list of words for what the model is good at, such as ["reasoning", "coding"], and shows in the TAGS column of agency local list. Both are optional. A kind that is not one of the four, or tags that are not a list of strings, is an error when the config is read. A catalog you host through modelCatalogUrl uses the same two fields.

See also ​

Chat with images ​

agency local serve --vlm qwen3.5-9b-mlx starts the model with mlx-vlm==0.7.0. The flag is repeatable and required for image-input chat. Without it, the model uses the text runtime. See chat with pictures for installation and an Agency example.

A shared serve option such as --hedge-limit is refused when no model in the command supports it. It is allowed when a text model in the same command uses it. --draft applies to the preceding model and is refused for a model named with --vlm.