Skip to content

vision ​

Ask a local vision model about an image: which objects are where, which booru tags describe it, a caption, or embeddings that say how parts of it look. The model must be running: start it with agency local serve <model>. The image never leaves the machine.

ts
import { detectObjects, tagImage } from "std::vision"
import { cropImage } from "std::image"

node main() {
  handle {
    const found = detectObjects("page.png", ["person", "desk", "chair"], "florence-2") catch []
    for (hit in found) {
      cropImage("page.png", hit.box, "crops/${hit.label}_${hit.id}.png", pad: 0.05)
    }
    const tags = tagImage("crops/person_0.png", "wd14-tagger") catch []
    const names = map(tags) as t {
      return t.tag
    }
    print(names.join(", "))
  } with approve
}

Each function is one job with a Result, so a labeling loop is a few lines of your own code over them, and partial application narrows each: detectObjects.partial(labels: ["person"], model: "florence-2") is a person-finder.

Finding your own things ​

A detector finds things by the name of their kind. It finds every mug, and cannot tell yours from the one next to it. To find your own mug, or the cats you draw, compare instead:

  1. findRegions boxes every thing in a picture, with no names. Where a detector knows the kind, detectObjects gives the boxes instead.
  2. embedImage turns each box into an embedding: a list of numbers that says how it looks. Things that look alike get embeddings that are close together.
  3. Compare each box with embeddings of crops of your own things, and of other things, and keep the boxes nearest yours.
ts
import { findRegions, embedImage } from "std::vision"
import { cosineSimilarity } from "std::embedding"

// How alike `vector` is to the closest of `references`, where 1 is identical.
def closest(vector: number[], references: number[][]): number {
  let best = 0
  for (reference in references) {
    const similarity = cosineSimilarity(vector, reference)
    if (isSuccess(similarity) && similarity.value > best) {
      best = similarity.value
    }
  }
  return best
}

// One vector per crop file.
def embedCrops(paths: string[]): number[][] {
  let vectors: number[][] = []
  for (path in paths) {
    const embedded = embedImage(path, "dinov2-base")
    if (isSuccess(embedded)) {
      vectors.push(embedded.value[0])
    }
  }
  return vectors
}

node main(page: string, catCrops: string[], otherCrops: string[]) {
  const cats = embedCrops(catCrops)
  const others = embedCrops(otherCrops)
  const regions = findRegions(page, "owlv2-base")
  if (isFailure(regions)) {
    return regions.error
  }
  const boxes = map(regions.value) as region {
    return region.box
  }
  const vectors = embedImage(page, "dinov2-base", boxes: boxes)
  if (isFailure(vectors)) {
    return vectors.error
  }
  for (vector, i in vectors.value) {
    if (closest(vector, cats) > closest(vector, others)) {
      print("Cat in region ${i}")
    }
  }
}

The crops are tight crops made with cropImage, one file each. A crop file embedded whole is cut the same way as a box embedded from a page, so the two compare fairly. A whole photo of a mug on a desk and a tight crop of a mug do not: the first vector mostly describes the desk.

Types ​

BoundingBox ​

Normalized to 0..1 with the origin at the top left of the image, so y grows downward. The same shape std::ocr returns, so a box from either can go to cropImage.

ts
/** Normalized to 0..1 with the origin at the top left of the image, so
`y` grows downward. The same shape std::ocr returns, so a box from
either can go to cropImage. */
export type BoundingBox = {
  x: number;
  y: number;
  width: number;
  height: number
}

(source)

Detection ​

One thing a detector found. id is its index in the reply, so crops can be named after it. The box is normalized 0..1 from the top left.

ts
/** One thing a detector found. `id` is its index in the reply, so crops
can be named after it. The box is normalized 0..1 from the top left. */
export type Detection = {
  id: number;
  label: string;
  score: number;
  box: BoundingBox
}

(source)

Region ​

One box around a thing in a picture, found with no name. id is its index in the reply, so crops can be named after it. The score says how much it looks like a thing at all, 0..1.

ts
/** One box around a thing in a picture, found with no name. `id` is its
index in the reply, so crops can be named after it. The score says how
much it looks like a thing at all, 0..1. */
export type Region = {
  id: number;
  score: number;
  box: BoundingBox
}

(source)

Tag ​

One tag a tagger gave, with how sure it was, 0..1.

ts
/** One tag a tagger gave, with how sure it was, 0..1. */
export type Tag = {
  tag: string;
  score: number
}

(source)

Effects ​

std::vision ​

ts
@alwaysUnder(dir)
effect std::vision {
  dir: string;
  filename: string;
  task: string;
  model: string
}

(source)

Functions ​

detectObjects ​

ts
detectObjects(
  path: string,
  labels: string[],
  model: string,
  threshold: number | null = null,
): Result<Detection[]> raises <std::vision>

Find the named things in an image with a local detector and return each one's label, score, and bounding box. The box is normalized to 0..1 with the origin at the top left, the shape cropImage takes. Runs on this machine; nothing is uploaded.

@param path - Path to a PNG, JPEG, WebP, or GIF image @param labels - What to look for, in words: ["person", "desk", "chair"]. A detector with no labels finds whatever it likes, so at least one is required. OWLv2 looks for every label in one pass; Florence-2 looks for each in its own pass, so each label adds to the time @param model - A local vision model that detects, such as "owlv2-base" or "florence-2" @param threshold - Drop detections scoring below this, 0 to 1. Null uses the model's default: 0.1 for OWLv2, whose real objects often score 0.2 to 0.4. Florence-2 gives no scores: every box it finds scores 1, so the threshold drops nothing

Parameters:

NameTypeDefault
pathstring
labelsstring[]
modelstring
thresholdnumber | nullnull

Returns: Result<Detection[]>

Throws: std::vision

(source)

tagImage ​

ts
tagImage(
  path: string,
  model: string,
  threshold: number = 0.35,
  limit: number = 30,
): Result<Tag[]> raises <std::vision>

Describe an image as booru tags with a local tagger: "1girl, glasses, reading, book". Each tag comes with its score, best first. This is the vocabulary illustration models such as NoobAI-XL take in a prompt, so the tags can go straight into a caption file. Runs on this machine; nothing is uploaded.

@param path - Path to a PNG, JPEG, WebP, or GIF image @param model - A local vision model that tags, such as "wd14-tagger" or "florence-2" @param threshold - Drop tags scoring below this, 0 to 1 @param limit - At most this many tags, 1 to 500

Parameters:

NameTypeDefault
pathstring
modelstring
thresholdnumber0.35
limitnumber30

Returns: Result<Tag[]>

Throws: std::vision

(source)

captionImage ​

ts
captionImage(
  path: string,
  model: string,
  detail: string = "short",
): Result<string> raises <std::vision>

Write a sentence about an image with a local model. Runs on this machine; nothing is uploaded.

@param path - Path to a PNG, JPEG, WebP, or GIF image @param model - A local vision model that captions, such as "florence-2" @param detail - "short" for one sentence, "long" for a paragraph

Parameters:

NameTypeDefault
pathstring
modelstring
detailstring"short"

Returns: Result<string>

Throws: std::vision

(source)

embedImage ​

ts
embedImage(
  path: string,
  model: string,
  boxes: BoundingBox[] | null = null,
): Result<number[][]> raises <std::vision>

Turn an image into an embedding, a list of numbers that says how it looks, so things that look alike come out close together. Pass boxes to embed each box instead of the whole image, one embedding per box. Compare two embeddings only when both were cut the same way, such as a tight crop and a box around the same thing.

@param path - Path to a PNG, JPEG, WebP, or GIF image @param model - A local vision model that embeds, such as "dinov2-base" @param boxes - Boxes to embed, as detectObjects and findRegions return them, at most 100. Null embeds the whole image

Similarity between two embeddings is cosineSimilarity from std::embedding, where 1 is identical. As a starting point, on one photo of two cats and two TV remotes, the two cats scored 0.55 against each other, the two remotes 0.57, and a cat against a remote at most 0.26. The same object in two different photos has not been measured, and neither have pen-and-ink drawings. Rather than pick a cut-off, compare each box with crops of your own things and of other things, and keep the nearer side: the module doc comment shows how.

Parameters:

NameTypeDefault
pathstring
modelstring
boxesBoundingBox[] | nullnull

Returns: Result<number[][]>

Throws: std::vision

(source)

findRegions ​

ts
findRegions(
  path: string,
  model: string,
  limit: number = 50,
  threshold: number | null = null,
): Result<Region[]> raises <std::vision>

Box every thing in an image with a local model, with no names: for finding things a detector has no word for, such as the characters in a drawing. Each region comes with a score for how much it looks like a thing at all, best first. Runs on this machine; nothing is uploaded.

@param path - Path to a PNG, JPEG, WebP, or GIF image @param model - A local vision model that finds regions, such as "owlv2-base" or "florence-2" @param limit - At most this many regions, 1 to 100 @param threshold - Drop regions scoring below this, 0 to 1. Null uses 0.1, which drops the many boxes OWLv2 puts on nothing. Florence-2 gives no scores: every region scores 1, so the threshold drops nothing

Parameters:

NameTypeDefault
pathstring
modelstring
limitnumber50
thresholdnumber | nullnull

Returns: Result<Region[]>

Throws: std::vision

(source)