Skip to content

speech ​

Speak text aloud locally, synthesize and transcribe speech in the cloud, and record from the microphone.

ts
import { record, transcribe, speak, say } from "std::speech"

node main() {
  const audio = record()
  const text = transcribe(audio)
  say("You said: ${text}")             // local playback
  const mp3 = speak("Cloud voice: ${text}")  // cloud TTS -> file path
  const wav = speakLocal("Local voice: ${text}", "qwen3-tts-mlx")  // local TTS -> file path
}

Effects ​

std::say ​

ts
effect std::say {
  textLength: number
}

(source)

std::record ​

ts
effect std::record {}

(source)

std::transcribe ​

ts
effect std::transcribe {
  requestedProvider: string;
  configuredModel: string;
  filepath: string
}

(source)

std::synthesizeSpeech ​

ts
effect std::synthesizeSpeech {
  requestedProvider: string;
  configuredModel: string;
  textLength: number;
  instructions: string
}

(source)

std::localSpeech ​

ts
effect std::localSpeech {
  model: string;
  voice: string;
  textLength: number;
  outputFile: string;
  format: string
}

(source)

Functions ​

say ​

ts
say(
  text: string,
  voice: string = "",
  rate: number = 0,
  outputFile: string = "",
  allowedPaths: string[] = [],
)

Speak text aloud locally using the operating system's text-to-speech (macOS only). For cloud text-to-speech that returns an audio file, use speak() instead.

@param text - The text to speak @param voice - Voice name to use @param rate - Speaking rate in words per minute @param outputFile - When set, save the audio to this file instead of playing it @param allowedPaths - Only allow saving under these path prefixes

Parameters:

NameTypeDefault
textstring
voicestring""
ratenumber0
outputFilestring""
allowedPathsstring[][]

Throws: std::say

(source)

record ​

ts
record(
  outputFile: string = "",
  silenceTimeout: number = 2000,
  allowedPaths: string[] = [],
): string

Record audio from the microphone. Recording stops when the user presses Enter, or after the silence timeout elapses.

@param outputFile - File path to save audio to (auto-generated in the temp directory if empty) @param silenceTimeout - Silence before auto-stopping, in milliseconds; 0 disables silence detection so recording stops only on Enter @param allowedPaths - Only allow saving a non-empty outputFile under these path prefixes

silenceTimeout is in milliseconds, so you can pass Agency's unit literals: record(silenceTimeout: 3s), record(silenceTimeout: 500ms).

Ctrl-C, a race loss, or a time-guard abort stops an in-progress recording, which surfaces as an AgencyCancelledError.

An empty outputFile is auto-generated under the system temp directory and is not subject to the allowedPaths allow-list.

Parameters:

NameTypeDefault
outputFilestring""
silenceTimeoutnumber2000
allowedPathsstring[][]

Returns: string

Throws: std::record

(source)

transcribe ​

ts
transcribe(
  filepath: string,
  language: string = "",
  allowedPaths: string[] = [],
  model: string = "whisper-1",
  provider: string = "",
  prompt: string = "",
  timestampGranularity: string = "",
  apiKey: string = "",
): string

Transcribe an audio file to text using a cloud speech provider (OpenAI Whisper by default). Returns the transcript text; throws on failure.

@param filepath - Path to the audio file @param language - Language code (e.g. "en") for better accuracy @param allowedPaths - Only allow reading audio files under these path prefixes @param model - Transcription model (default "whisper-1") @param provider - Override the provider (normally derived from the model name) @param prompt - Optional text to bias decoding (names, jargon) @param timestampGranularity - "segment" or "word" to request timestamps @param apiKey - Override the API key

A cloud transcription request tears down on Ctrl-C, race-loser, or time-guard abort. Cost, spend guards, and statelog apply.

Parameters:

NameTypeDefault
filepathstring
languagestring""
allowedPathsstring[][]
modelstring"whisper-1"
providerstring""
promptstring""
timestampGranularitystring""
apiKeystring""

Returns: string

Throws: std::transcribe

(source)

speak ​

ts
speak(
  text: string,
  outputFile: string = "",
  voice: string = "alloy",
  model: string = "tts-1",
  provider: string = "",
  format: string = "mp3",
  speed: number = 1,
  allowedPaths: string[] = [],
  apiKey: string = "",
  instructions: string = "",
): string

Synthesize speech from text using a cloud text-to-speech provider (OpenAI by default), writing the audio to a file and returning its path. For local playback through the speakers, use say() instead.

@param text - The text to synthesize @param outputFile - Where to write the audio (auto-generated temp file if empty). Never overwrites an existing file. @param voice - Voice name (default "alloy") @param model - Speech model (default "tts-1") @param provider - Override the provider (normally derived from the model name) @param format - Output audio format: mp3 (default), opus, aac, flac, wav, or pcm @param speed - Speaking speed (0.25 to 4.0) @param allowedPaths - Only allow writing a non-empty outputFile under these path prefixes @param apiKey - Override the API key @param instructions - How the speech should sound, in plain words: "Calm and slow." Needs a model that reads it, such as gpt-4o-mini-tts; tts-1 and tts-1-hd ignore it

A cloud synthesis request tears down on Ctrl-C, race-loser, or time-guard abort; a cancelled request never writes its output file. Cost, spend guards, and statelog apply.

Parameters:

NameTypeDefault
textstring
outputFilestring""
voicestring"alloy"
modelstring"tts-1"
providerstring""
formatstring"mp3"
speednumber1
allowedPathsstring[][]
apiKeystring""
instructionsstring""

Returns: string

Throws: std::synthesizeSpeech

(source)

speakLocal ​

ts
speakLocal(
  text: string,
  model: string,
  voice: string = "",
  instructions: string = "",
  outputFile: string = "",
  format: string = "",
  allowedPaths: string[] = [],
  speed: number = 1,
): string raises <std::localSpeech>

Speak text into an audio file with a local speech model, and return the path of that file.

@param text - The text to speak. Orpheus models perform emotion tags written in the text, such as <laugh>, <sigh>, or <gasp>. @param model - The speech model: a catalog name such as "qwen3-tts-mlx" or "orpheus-3b-mlx", an mlx: URI, or a repo id the server is serving. @param voice - A voice of that model. Leave empty for the model's default. VoiceDesign models take no voice. @param instructions - How the speech should sound, such as "Alarmed and urgent". Qwen3-TTS models only; Orpheus takes tags in the text instead. @param outputFile - Where to write the file. Leave empty for a new temp file. An existing file is never overwritten. @param format - "wav", "mp3", "m4a", or "pcm". Leave empty to use the output file's extension, or wav when there is none. A format you give wins over the extension. @param allowedPaths - Directories that outputFile must be inside @param speed - Speaking speed, from 0.5 to 100 (both included). Anything other than 1 needs ffmpeg.

Use this for speech that must not leave the machine, and for the emotion a local model can perform that a cloud voice cannot. Long text is sent to the server in pieces, so Ctrl-C stops within one piece and writes no file. The usage record costs nothing, and statelog gets one event per call however many pieces the text becomes.

Writing mp3 or m4a, or a speed other than 1, needs ffmpeg on the PATH. speakLocal refuses such a call before raising any interrupt when ffmpeg is missing. The local models cannot change their own speed, so speed stretches the audio after it is generated, keeping its pitch.

Parameters:

NameTypeDefault
textstring
modelstring
voicestring""
instructionsstring""
outputFilestring""
formatstring""
allowedPathsstring[][]
speednumber1

Returns: string

Throws: std::localSpeech

(source)