First three pieces of the 0.8 multimodal epic ship today. The rest of the epic (provider adapters, KSP routing, manifest-anchored capability checks) is staged on top of this foundation.
sealed interface Content(#2466) —Text,Image,Audio,Video,Document. Each non-text variant carries aContentRef+ a typed mime (ImageMime,AudioMime,VideoMime,DocMime). NoStringmime anywhere in the public API.ContentRef+BlobStore(#2467) — content-addressed reference (SHA-256 hex + size + wire mime).InMemoryBlobStorefor tests,FileBlobStore(dir)for on-disk persistence that survives process restart.ToolResult(parts: List<Content>)(#2469) — tools can return mixed content. JSONL audit exporter records per-part modality + ref summary; no blob bytes ever enter the audit row.
- Modality + format live in the type, never in a String. Adding a new mime is a new variant; mistyping it is a compile error.
- Content-addressed payload, not inlined bytes.
Contentcarries aContentRef, blobs live in aBlobStore. Snapshot files stay small. Audit rows stay small. - No
data classholdingByteArray. Kotlin data-class equals/hashCode treat arrays by identity, not content — would break every downstream consumer's assumption. The ref pattern sidesteps it entirely.
import agents_engine.content.*
val store = FileBlobStore(Path.of("snapshots", "blobs"))
val screenshotAgent = agent<String, String>("screenshot") {
model { ollama("test") }
tools {
tool("capture", "Capture a screenshot of the URL") { args ->
val bytes = captureBytes(args["url"] as String)
val ref = store.put(bytes, ImageMime.Png.wireMime)
ToolResult(
Content.Text("Captured ${ref.sizeBytes}B image."),
Content.Image(ref, ImageMime.Png),
)
}
}
skills { skill<String, String>("respond", "") { tools("capture") } }
}sealed interface Content {
data class Text(val text: String) : Content
data class Image(val ref: ContentRef, val mime: ImageMime) : Content
data class Audio(val ref: ContentRef, val mime: AudioMime) : Content
data class Video(val ref: ContentRef, val mime: VideoMime) : Content
data class Document(val ref: ContentRef, val mime: DocMime) : Content
}
val Content.modality: String // "text" | "image" | "audio" | "video" | "document"Stage 1 (this release): Image wire path (provider input + audit + tool-result rendering) is live. Document flows through the audit/ToolResult placeholder path but the provider-input adapter (#2470 slice c) is deferred — agent.invokeWithAttachments(...) silently skips Content.Document today. Audio is end-to-end via the speech tools (#4501 — transcribe_audio / speak; see "End-to-end audio" below) rather than the chat-message attachment path; Video is modelled but not yet wired. See providers.md for the per-provider × content-type support matrix — the single source of truth for what's routed to provider input today.
sealed interface ImageMime { Png, Jpeg, Gif, Webp }
sealed interface AudioMime { Mp3, Wav, Flac, Ogg }
sealed interface VideoMime { Mp4, WebM, Mov }
sealed interface DocMime { Pdf, Docx, Markdown, Html, PlainText }Each variant exposes wireMime: String for adapter serialisation. The public API never accepts String mime — extending is adding a variant.
The boilerplate "read file, put bytes, build Content variant" pattern compresses to one call:
import agents_engine.content.Files
agent.invokeWithAttachments(
"Review this spec",
attachments = listOf(Files.load(Path.of("spec.pdf"), store)),
)Files.load(path, store) reads the file, detects modality + mime from the filename extension, puts the bytes into the BlobStore, and returns the right typed Content variant. Same ContentRef.hash as calling store.put(bytes, mime) directly.
Extension coverage:
| Modality | Extensions |
|---|---|
| Image | png, jpg, jpeg, gif, webp |
| Audio | mp3, wav, flac, ogg |
| Video | mp4, webm, mov |
| Document | pdf, docx, md, markdown, html, htm, txt |
Case-insensitive (Foo.PDF and foo.pdf both produce DocMime.Pdf). Detection is by extension only — no magic-byte sniffing. Fast, predictable; if the extension lies, build Content.X(ref, mime) explicitly.
Variants:
Files.load(path, store, maxBytes = 20.MB): Content // throws on unknown ext or oversize
Files.loadOrNull(path, store, maxBytes = 20.MB): Content? // null on unknown; throws on oversize
Files.loadAll(paths, store, maxBytes = 20.MB): List<Content>
Files.loadAllOrSkip(paths, store, maxBytes = 20.MB) // skip unknown ext; oversize still throws
Files.canonicalExtensionFor(content): String? // inverse — write Content back to disk
Files.knownExtensions: Set<String> // "can I load this?" predicate
Files.DEFAULT_MAX_BYTES // 20 MiBUnknownExtensionException names the offending extension, the path, AND lists every known extension — debuggable on a misconfigured callsite. OversizedFileException (#2871) names the path, actual size, and configured cap; the size check uses Files.size(path) before any bytes are read, so an accidental 4 GiB upload fails-fast without OOMing the JVM. BlobStore.verify(ref) re-hashes stored bytes and returns false on corruption or absence — opt-in integrity check, not on the hot path of get.
data class ContentRef(val hash: String, val sizeBytes: Long, val wireMime: String)
interface BlobStore {
fun put(bytes: ByteArray, wireMime: String): ContentRef
fun get(ref: ContentRef): ByteArray?
fun open(ref: ContentRef): InputStream?
fun exists(ref: ContentRef): Boolean
fun delete(ref: ContentRef)
}
class InMemoryBlobStore : BlobStore
class FileBlobStore(dir: Path) : BlobStoreHash: SHA-256 hex. Same family as manifest hash (#1912) and snapshot filename hash (#2753) — single hash algorithm across the audit surface.
Idempotent put: putting the same bytes twice returns the same ContentRef. FileBlobStore writes the file once and is a no-op on the second put — pinned by a test.
Persistence: FileBlobStore survives process restart. A fresh instance on the same directory sees blobs from prior puts. Atomic via tmp + rename, matching the FileSnapshotStore pattern (#2753).
Custom backends: an internal artifact registry, S3, GCS, etc. implement BlobStore and plug in via the same interface.
tool("ocr", "Extract text + return source") { args ->
val text = ocrText(args["pdf"] as String)
val sourceRef = store.put(args["pdf-bytes"] as ByteArray, DocMime.Pdf.wireMime)
ToolResult(
Content.Text(text),
Content.Document(sourceRef, DocMime.Pdf),
)
}Tools can return ToolResult instead of a String. The agentic loop renders it for the LLM's tool-result message in v1:
Extracted spec text.
[document: application/pdf] (a3f9b2c4...12, 524288B)
Provider-specific multipart rendering — the model actually seeing the image / document — lands in #2470 (Provider normalization adapters, deferred).
untrustedOutput = true still wraps the rendered text summary in the JSON envelope (#642 composes with #2469).
JsonlAuditExporter (#1914) gains a new outputParts column. For ToolResult returns:
{
"...": "...",
"outputType": "agents_engine.content.ToolResult",
"outputParts": [
"text:inline:18:text/plain",
"image:a3f9b2c41205:524288:image/png"
],
"...": "..."
}Format: <modality>:<hash-prefix>:<sizeBytes>:<wireMime> per part. Hash prefix is the first 12 hex chars — enough to disambiguate in audit grep, short enough to keep audit rows compact.
Critical: blob bytes never enter the audit row. The ContentRef is the auditable surface. Pinned by a dedicated test that asserts the PNG magic byte sequence does not appear anywhere in the audit JSON.
The same discipline applies (when wired) to the OTel / LangSmith / Langfuse bridges — sibling tickets will plumb outputParts onto span events / run events / observations.
Content carrying a ContentRef (not bytes) means SessionSnapshot (#2386 / #2754) stays small regardless of how much image / audio / video flowed through the agent. The snapshot serialises the ref; the blob lives in the BlobStore. A snapshot resume against the same BlobStore dereferences the ref normally.
Pairs with the #2754 manifest-hash restore guard: resume across an agent rebuild that changed tools (including the BlobStore wiring) fails closed unless the caller opts in.
The foundation above answers "how do tools return mixed content?" The follow-on slice answers "how does the agent send an image to a vision-capable model?"
LlmMessage gains an optional images: List<ImagePart> field. Adapters translate it per provider:
import agents_engine.model.LlmMessage
import agents_engine.model.ImagePart
val png: ByteArray = Files.readAllBytes(Path.of("screenshot.png"))
val client = OllamaClient(model = "qwen3-vl:8b", temperature = 0.0)
client.chat(listOf(
LlmMessage(
role = "user",
content = "How many windows in this picture?",
images = listOf(ImagePart(
base64 = Base64.getEncoder().encodeToString(png),
wireMime = ImagePart.WireMime.Png,
)),
),
))ImagePart carries the base64 payload + a closed WireMime sealed type (Png, Jpeg, Gif, Webp). String mime is not accepted in the public ctor — same closed-mime discipline as Content.Image. Caller base64-encodes upfront so the adapter can splat the payload straight onto the wire without re-encoding per provider.
| Provider | User-message shape |
|---|---|
| Ollama | {role:"user", content:"text", images:["<b64>", ...]} |
| Claude | {role:"user", content:[{type:"text"}, {type:"image", source:{type:"base64", media_type:"image/png", data:"<b64>"}}, ...]} |
| OpenAI | {role:"user", content:[{type:"text"}, {type:"image_url", image_url:{url:"data:image/png;base64,<b64>"}}, ...]} |
| DeepSeek | inherits OpenAI; current DeepSeek models lack vision and silently ignore the field |
Back-compat: images = null (or empty) produces byte-identical wire shape to pre-#2470. Pinned by dedicated tests.
Role-gating: images are emitted only on role = "user" messages. System / assistant / tool messages with non-null images ignore the field on the wire — no provider's API carries images on those roles.
VisionFixtures (in src/test) ships two reproducible PNGs rendered via BufferedImage:
threeSquaresPng()— 256×256, three colored squares (red/blue/green) on white with thick black outlines. Used for "count the squares" eval against cheap vision models.housePng()— 256×256 cartoon house: triangle roof + body + door + two windows. Used for "what is this?" eval.
No external assets, no binary blobs in the repo — every byte is reproducible across machines and CI runs.
VisionLiveTest.kt runs the two fixtures against all three vision-capable providers, with cost discipline (256×256 PNG ~5KB, temperature = 0, maxTokens = 80, single-turn):
| Provider | Default model | How to run |
|---|---|---|
| Ollama | qwen3-vl:8b |
./gradlew integrationTest --tests "*VisionLiveTest*" (tagged live-llm) |
| Claude | claude-haiku-4-5 |
./gradlew test --tests "*VisionLiveTest*" (tagged live-cloud-api; assumeTrue skips when no key) |
| OpenAI | gpt-4o-mini |
same |
Models overridable via env (AGENTSKT_TEST_OLLAMA_VISION_MODEL, AGENTSKT_TEST_CLAUDE_VISION_MODEL, AGENTSKT_TEST_OPENAI_VISION_MODEL).
Assertion shape is loose: the test passes if the model's text response mentions one of a small acceptable keyword set (3 / three for counting; house / home / cottage / building / cabin / barn for the house). Goal is "did the image reach the model and elicit a sensible reply" — not exact phrasing.
Slice a wires the per-provider wire format on LlmMessage.images. Slice b puts a clean user-facing API on top: the caller passes typed Content.Image (carrying a ContentRef) at the agent's invoke surface; the runtime dereferences against an injected BlobStore, base64-encodes once, and attaches ImagePart to the first user message.
import agents_engine.content.Content
import agents_engine.content.ImageMime
import agents_engine.content.FileBlobStore
val store = FileBlobStore(Path.of("snapshots/blobs"))
val agent = agent<String, String>("vision") {
model { ollama("qwen3-vl:8b") }
blobStore(store)
skills { skill<String, String>("describe", "") { tools() } }
}
val ref = store.put(pngBytes, ImageMime.Png.wireMime)
val reply = agent.invokeWithAttachments(
"What is in this image?",
attachments = listOf(Content.Image(ref, ImageMime.Png)),
)- First user message only. Attachments ride on the initial user turn. Multi-turn vision is composable but each invocation owns its own first-turn attachments.
- Closed-mime mapping.
ImageMime → ImagePart.WireMimefor all four variants —Png,Jpeg,Gif,Webp. No String conversions anywhere. - Fail-fast on misconfiguration. Passing attachments to an agent with no
blobStoreconfigured errors fast at invoke time withAgent '<name>' has attachments but no blobStore. A ref pointing at a missing blob errors fast with the ref's hash prefix in the message for forensics. - Skip non-image variants.
Content.Text/Content.Document/Content.Audio/Content.Videoare silently skipped in v1 — slice c will wire Document via provider doc-input adapters, audio/video as part of Stage 2. - Back-compat.
agent.invokeSuspend(input)(without attachments) stays byte-identical on the wire. The attachments path is purely additive — opt in via the newinvokeWithAttachments/invokeSuspendWithAttachmentsentry points. - Snapshot/resume composition. On resume the saved user turn already carries the original
LlmMessage.images; the runtime ignores theattachmentsargument because the conversation was restored intact.
| Entry point | Use when |
|---|---|
agent.invokeSuspendWithAttachments(input, attachments) |
Inside coroutine scopes — composition operators, structured concurrency. |
agent.invokeWithAttachments(input, attachments) |
Outside coroutine scopes — quick scripts, REPL, blocking glue. Thin runBlocking shim. |
Mirrors the existing invokeSuspend / invoke split.
AgentVisionLiveTest.kt runs the two VisionFixtures (threeSquaresPng() + housePng()) through the agent surface on all three vision-capable providers. Same cost discipline as slice a — 256×256 PNG, temperature = 0, maxTokens = 80, single-turn. Tagged live-llm (Ollama) / live-cloud-api (Claude + OpenAI); assumeTrue skips per-provider when no key.
The slice-b live tests complement the slice-a tests: the slice-a VisionLiveTest exercises the raw ModelClient; slice-b's AgentVisionLiveTest exercises the full agent loop including BlobStore deref.
- #2468 Compile-time modality routing —
Agent<Image, X>becomes a real type; cross-modality miswiring is a compile error. Multi-part@Generableinputs via KSP. - #2470 slice c Document/Audio/Video provider-input adapters — currently only images flow through the wire; Document/Audio/Video Content variants are skipped on the attachment path.
- #2471 Manifest-anchored modality capability — declared per-agent modalities recorded in the permission manifest, validated against provider capabilities at build time.
- #2472 Multimodal memory —
MemoryBankentries carryContentReffor image/audio/video state. - #2473 Testing fixtures + snapshot + mutation coverage.
Stage 2 (Audio + Video) lights up when a concrete use case lands.
docs/permission-manifest.md— the manifest-hash family the BlobStore reuses.docs/observability.md— the audit exporter that now carriesoutputParts.docs/hitl.md—Contentwill appear in human-approval bodies once Stage 2 lands.
Sources: agents_engine/content/Content.kt, agents_engine/content/ContentRef.kt, agents_engine/content/ToolResult.kt, audit wiring in agents-kt-observability/.../JsonlAuditExporter.kt.
Tests: ContentAndRefTest.kt, ToolResultIntegrationTest.kt, JsonlAuditExporterTest's "multimodal ToolResult writes outputParts" case.
Three typed clients close the audio/image-generation gap, each storing bytes in your BlobStore and passing the typed ref through the agent graph:
val transcript: String = OpenAiSpeechToTextClient(key).transcribe(audioContent, blobStore) // Whisper
val image: Content.Image = OpenAiImagesClient(key, blobStore).generate(transcript) // Images API
val speech: Content.Audio = OpenAiTtsClient(key, blobStore).speak("done — image attached") // TTSSpeechToTextClient / ImageModelClient / TtsModelClient are fun interfaces — bring another provider by implementing one method. Wire shapes are pinned by stub-server tests; baseUrl is injectable. Not yet wired: Content.Audio directly inside chat messages (gpt-4o-audio-style input_audio blocks) — transcribe-then-prompt is the v1 path, matching the roadmap's STT-helper framing.
The standalone clients above let you call STT/TTS by hand. To make audio end-to-end through an agent, speech is also exposed as tools — the model decides when to transcribe or speak, and the calls ride the full tool spine (ToolPolicy + the Layer-1 filesystem gate, constraints, audit, manifest, typed hooks). No attachment-path wiring needed: ToolResult already carries Content.Audio, so the speak tool's output ref flows through audit and snapshot like any other content.
val blobs = InMemoryBlobStore()
val agent = agent<String, String>("voicebot") {
model { ollama("llama3.1") }
tools {
+transcribeAudioTool(WhisperSttClient("http://localhost:8000"), blobs, audioRoot = "/recordings")
+speakTool(QwenTtsClient("http://localhost:8880", blobs, voice = "Cherry"))
}
skills { skill<String, String>("assist", "Hears and answers") { tools("transcribe_audio", "speak") } }
}transcribe_audio(path)— loads a local audio file (confined toaudioRootby a declaredfilesystem.readpolicy and a defensive prefix check), runs STT, returns the transcript. Declaresnetworkegress (the client ships bytes to its endpoint).speak(text)— runs TTS and returns aToolResult(Content.Text, Content.Audio); the synthesized bytes live in theBlobStore, the typed ref travels.speechTools(stt, tts, blobs, audioRoot)bundles both.
Both target the OpenAI-compatible wire (/v1/audio/transcriptions, /v1/audio/speech) that the common self-hosted servers expose, with no API key by default and a required baseUrl:
WhisperSttClient(baseUrl = "http://localhost:8000") // faster-whisper-server / Speaches / LocalAI
QwenTtsClient(baseUrl = "http://localhost:8880", blobStore = blobs, voice = "Cherry") // openedai-speech / LocalAIBoth implement the existing SpeechToTextClient / TtsModelClient interfaces, so the OpenAI hosted adapters remain drop-in swappable in the same tools. Pass bearerToken only when a gateway in front of the server requires one. QwenTtsClient's responseFormat (mp3/wav/flac/opus) maps to the typed AudioMime.
Fail-fast DX (#4504). Self-hosted endpoints are routinely "not running yet", so call client.preflight() at startup — it probes the endpoint and throws an actionable message ("start a server … or fix baseUrl") instead of failing later with a raw ConnectException. transcribe/speak wrap connection/timeout failures into the same actionable error.
val stt = WhisperSttClient("http://localhost:8000").apply { preflight() } // fails loud now if the server is downagents.kt is an orchestration toolkit, not a media server — it never binds a listening port. The HTTP clients above are outbound (they connect to a service you run; nothing is exposed on our side). When you want the engine on the same box with no server at all, bolt it on through the same seams via one of two portless transports:
- In-process (JNI) —
:agents-kt-whisper-jni'sWhisperJniSttClientruns whisper.cpp inside the JVM. ASpeechToTextClient, drop-in fortranscribe_audio. No port. - Subprocess —
SubprocessTtsClientpipes text to a local TTS binary (e.g.piper) throughProcessSandbox(stdin in, audio file out), write-confined to a temp dir, and returnsContent.Audio. ATtsModelClient, drop-in forspeak. No port, no JVM TTS engine — the engine is the external binary.
val tts = SubprocessTtsClient(blobStore) { _, out ->
listOf("piper", "--model", "/voices/en.onnx", "--output_file", out.toString())
}
agent { tools { +speakTool(tts) } } // local TTS, engine external, zero portsSo voice has three shapes, all behind the same SpeechToTextClient / TtsModelClient seams: outbound HTTP (a service you run), in-process (JNI), or subprocess (a local binary). Pick by where the engine lives; the agent code is identical.
Agents.KT ships no model weights — not Whisper's, not Qwen-TTS's. The jar is code: the HTTP adapters above talk to a server you run (weights live in that server's process), so there is nothing to bundle. This keeps the artifact small and keeps weight licensing (Qwen license, model-card terms) entirely separate from the Apache-licensed code. Weights are a deploy-time concern resolved by baseUrl (HTTP) — or, for the in-process path, a model file you provision (:agents-kt-whisper-jni, which downloads + checksums a model at runtime, still bundling nothing).