vision
Ask a local vision model about an image: which objects are where, which booru tags describe it, a caption, or embeddings that say how parts of it look. The model must be running: start it with agency local serve <model>. The image never leaves the machine.
import { detectObjects, tagImage } from "std::vision"
import { cropImage } from "std::image"
node main() {
handle {
const found = detectObjects("page.png", ["person", "desk", "chair"], "florence-2") catch []
for (hit in found) {
cropImage("page.png", hit.box, "crops/${hit.label}_${hit.id}.png", pad: 0.05)
}
const tags = tagImage("crops/person_0.png", "wd14-tagger") catch []
const names = map(tags) as t {
return t.tag
}
print(names.join(", "))
} with approve
}Each function is one job with a Result, so a labeling loop is a few lines of your own code over them, and partial application narrows each: detectObjects.partial(labels: ["person"], model: "florence-2") is a person-finder.
Finding your own things
A detector finds things by the name of their kind. It finds every mug, and cannot tell yours from the one next to it. To find your own mug, or the cats you draw, compare instead:
findRegionsboxes every thing in a picture, with no names. Where a detector knows the kind,detectObjectsgives the boxes instead.embedImageturns each box into an embedding: a list of numbers that says how it looks. Things that look alike get embeddings that are close together.- Compare each box with embeddings of crops of your own things, and of other things, and keep the boxes nearest yours.
import { findRegions, embedImage } from "std::vision"
import { cosineSimilarity } from "std::embedding"
// How alike `vector` is to the closest of `references`, where 1 is identical.
def closest(vector: number[], references: number[][]): number {
let best = 0
for (reference in references) {
const similarity = cosineSimilarity(vector, reference)
if (isSuccess(similarity) && similarity.value > best) {
best = similarity.value
}
}
return best
}
// One vector per crop file.
def embedCrops(paths: string[]): number[][] {
let vectors: number[][] = []
for (path in paths) {
const embedded = embedImage(path, "dinov2-base")
if (isSuccess(embedded)) {
vectors.push(embedded.value[0])
}
}
return vectors
}
node main(page: string, catCrops: string[], otherCrops: string[]) {
const cats = embedCrops(catCrops)
const others = embedCrops(otherCrops)
const regions = findRegions(page, "owlv2-base")
if (isFailure(regions)) {
return regions.error
}
const boxes = map(regions.value) as region {
return region.box
}
const vectors = embedImage(page, "dinov2-base", boxes: boxes)
if (isFailure(vectors)) {
return vectors.error
}
for (vector, i in vectors.value) {
if (closest(vector, cats) > closest(vector, others)) {
print("Cat in region ${i}")
}
}
}The crops are tight crops made with cropImage, one file each. A crop file embedded whole is cut the same way as a box embedded from a page, so the two compare fairly. A whole photo of a mug on a desk and a tight crop of a mug do not: the first vector mostly describes the desk.
Types
BoundingBox
Normalized to 0..1 with the origin at the top left of the image, so y grows downward. The same shape std::ocr returns, so a box from either can go to cropImage.
/** Normalized to 0..1 with the origin at the top left of the image, so
`y` grows downward. The same shape std::ocr returns, so a box from
either can go to cropImage. */
export type BoundingBox = {
x: number;
y: number;
width: number;
height: number
}(source)
Detection
One thing a detector found. id is its index in the reply, so crops can be named after it. The box is normalized 0..1 from the top left.
/** One thing a detector found. `id` is its index in the reply, so crops
can be named after it. The box is normalized 0..1 from the top left. */
export type Detection = {
id: number;
label: string;
score: number;
box: BoundingBox
}(source)
Region
One box around a thing in a picture, found with no name. id is its index in the reply, so crops can be named after it. The score says how much it looks like a thing at all, 0..1.
/** One box around a thing in a picture, found with no name. `id` is its
index in the reply, so crops can be named after it. The score says how
much it looks like a thing at all, 0..1. */
export type Region = {
id: number;
score: number;
box: BoundingBox
}(source)
Tag
One tag a tagger gave, with how sure it was, 0..1.
/** One tag a tagger gave, with how sure it was, 0..1. */
export type Tag = {
tag: string;
score: number
}(source)
Effects
std::vision
@alwaysUnder(dir)
effect std::vision {
dir: string;
filename: string;
task: string;
model: string
}(source)
Functions
detectObjects
detectObjects(
path: string,
labels: string[],
model: string,
threshold: number | null = null,
): Result<Detection[]> raises <std::vision>Find the named things in an image with a local detector and return each one's label, score, and bounding box. The box is normalized to 0..1 with the origin at the top left, the shape cropImage takes. Runs on this machine; nothing is uploaded.
@param path - Path to a PNG, JPEG, WebP, or GIF image @param labels - What to look for, in words: ["person", "desk", "chair"]. A detector with no labels finds whatever it likes, so at least one is required. OWLv2 looks for every label in one pass; Florence-2 looks for each in its own pass, so each label adds to the time @param model - A local vision model that detects, such as "owlv2-base" or "florence-2" @param threshold - Drop detections scoring below this, 0 to 1. Null uses the model's default: 0.1 for OWLv2, whose real objects often score 0.2 to 0.4. Florence-2 gives no scores: every box it finds scores 1, so the threshold drops nothing
Parameters:
| Name | Type | Default |
|---|---|---|
| path | string | |
| labels | string[] | |
| model | string | |
| threshold | number | null | null |
Returns: Result<Detection[]>
Throws: std::vision
(source)
tagImage
tagImage(
path: string,
model: string,
threshold: number = 0.35,
limit: number = 30,
): Result<Tag[]> raises <std::vision>Describe an image as booru tags with a local tagger: "1girl, glasses, reading, book". Each tag comes with its score, best first. This is the vocabulary illustration models such as NoobAI-XL take in a prompt, so the tags can go straight into a caption file. Runs on this machine; nothing is uploaded.
@param path - Path to a PNG, JPEG, WebP, or GIF image @param model - A local vision model that tags, such as "wd14-tagger" or "florence-2" @param threshold - Drop tags scoring below this, 0 to 1 @param limit - At most this many tags, 1 to 500
Parameters:
| Name | Type | Default |
|---|---|---|
| path | string | |
| model | string | |
| threshold | number | 0.35 |
| limit | number | 30 |
Returns: Result<Tag[]>
Throws: std::vision
(source)
captionImage
captionImage(
path: string,
model: string,
detail: string = "short",
): Result<string> raises <std::vision>Write a sentence about an image with a local model. Runs on this machine; nothing is uploaded.
@param path - Path to a PNG, JPEG, WebP, or GIF image @param model - A local vision model that captions, such as "florence-2" @param detail - "short" for one sentence, "long" for a paragraph
Parameters:
| Name | Type | Default |
|---|---|---|
| path | string | |
| model | string | |
| detail | string | "short" |
Returns: Result<string>
Throws: std::vision
(source)
embedImage
embedImage(
path: string,
model: string,
boxes: BoundingBox[] | null = null,
): Result<number[][]> raises <std::vision>Turn an image into an embedding, a list of numbers that says how it looks, so things that look alike come out close together. Pass boxes to embed each box instead of the whole image, one embedding per box. Compare two embeddings only when both were cut the same way, such as a tight crop and a box around the same thing.
@param path - Path to a PNG, JPEG, WebP, or GIF image @param model - A local vision model that embeds, such as "dinov2-base" @param boxes - Boxes to embed, as detectObjects and findRegions return them, at most 100. Null embeds the whole image
Similarity between two embeddings is cosineSimilarity from std::embedding, where 1 is identical. As a starting point, on one photo of two cats and two TV remotes, the two cats scored 0.55 against each other, the two remotes 0.57, and a cat against a remote at most 0.26. The same object in two different photos has not been measured, and neither have pen-and-ink drawings. Rather than pick a cut-off, compare each box with crops of your own things and of other things, and keep the nearer side: the module doc comment shows how.
Parameters:
| Name | Type | Default |
|---|---|---|
| path | string | |
| model | string | |
| boxes | BoundingBox[] | null | null |
Returns: Result<number[][]>
Throws: std::vision
(source)
findRegions
findRegions(
path: string,
model: string,
limit: number = 50,
threshold: number | null = null,
): Result<Region[]> raises <std::vision>Box every thing in an image with a local model, with no names: for finding things a detector has no word for, such as the characters in a drawing. Each region comes with a score for how much it looks like a thing at all, best first. Runs on this machine; nothing is uploaded.
@param path - Path to a PNG, JPEG, WebP, or GIF image @param model - A local vision model that finds regions, such as "owlv2-base" or "florence-2" @param limit - At most this many regions, 1 to 100 @param threshold - Drop regions scoring below this, 0 to 1. Null uses 0.1, which drops the many boxes OWLv2 puts on nothing. Florence-2 gives no scores: every region scores 1, so the threshold drops nothing
Parameters:
| Name | Type | Default |
|---|---|---|
| path | string | |
| model | string | |
| limit | number | 50 |
| threshold | number | null | null |
Returns: Result<Region[]>
Throws: std::vision
(source)