LabsAI · Research Brief · Model Families

Introducing Actras & Octras

The Labs model families for speech that acts, and orchestration that listens.

August 15, 2026

An Actra converts speech and world state into action. An Octra converts objective and resources into orchestration policy. They are the model families behind Speech‑to‑Action (STA) and Speech‑to‑Orchestration (STO) — the two competences a spoken system needs the moment it stops merely talking and starts doing — and they are presented here the way we hold ourselves to presenting anything: with the interface that specifies them live in production, the benchmark that scores them published with its rubric, and every number a recorded run.

Labsintelligence · lab1 of Labs Actras · Octras · the LAIMA ladder

The claim underneath everything on this page. A model class is defined by its native prediction target, not by whether it contains text. Language models predict tokens. Actras predict action — sequences, parameters, state transitions, tool operations, confirmations — with language as one modality among several. Octras predict coordination — which intelligence does the work, how it is decomposed, when it escalates, how it is verified. The argument is made in full in Larger Than Language; this page is the family presentation.

01Two families

Actras

Action-Centered Transducers for Reasoning and Agentic Systems · singular: an Actra

Implements the STA class — Speech-to-Action

An Actra takes speech, conversation, application state, memory, intent and the vocabulary of available actions, and produces action: what the surface does, co-timed with what the voice says. The name is a transducer's name — it describes the conversion performed, not the topology that performs it — because the family is architecture-agnostic by design. Actras are to speech what VLA models are to vision — and in the terms that made the last decade legible, an Actra is a GAT: a Generative Action Transducer, to action what the generative pre-trained transformer was to language.

Octras

Orchestration-Centered Transducers for Routing and Agentic Systems · singular: an Octra

Implements the Orch class — Speech-to-Orchestration

An Octra takes state, objective and resources, and produces orchestration policy: model and tool selection, decomposition, parallel versus sequential execution, escalation, verification paths, termination — and in the spoken loop, all of it steered live by speech. STA decides what the surface does; STO decides which intelligence does it. The same move again: an Octra is a GOT — a Generative Orchestration Transducer, generating the coordination itself rather than the words about it. What it generates runs as an AI Orch — the executable coordination object that binds models, agents, tools, workflows, memory, policies, data and compute into one live operation. Orch Models are intelligence; AI Orchs are execution structures.

The two families are a designed pair: Action against Orchestration, Reasoning against Routing — the expansions differ exactly where the competences differ. Spoken, the two families’ work arrives as voice commands — Speech‑to‑Action Voice Commands for an Actra, Speech‑to‑Orchestration Voice Commands for an Octra — the spoken units the classes interpret. The shortest form of the whole stack: models reason, agents act, Orchs coordinate — a model thinks, an agent acts, a workflow follows predefined steps, and an AI Orch dynamically coordinates all of the above. What ships in production today is the first-generation Actra surface and Octra layer, live in LabsAI Studio; the composition of the shipped system is not published, as a standing policy stated once rather than repeated — read every claim on this page as a claim about the system as composed, measured from outside, the way its visitors meet it.

02What the first generation does

Builds the interface while speaking

Charts, tables, cards, playing video and live web pages are composed into a shared visual stage during the sentence that explains them — generative UI co-timed with narration, not attached after it. Arriving early is a slideshow; arriving late is a footnote; the bench scores the offset in milliseconds, signed.

Takes spoken commands against the live surface

Every element on the stage publishes its own voice-addressable affordances for exactly as long as it exists. "Pause it", "fold the second one", "no — the other one": referents resolve to specific on-screen objects, ordinals and corrections included, and a genuinely ambiguous reference is met with a question rather than a guess.

Verifies before it reports

An action is verified, attempted, blocked or failed — checked against the real interface, not assumed — and the report is bound to the outcome. A destructive action is held for confirmation. A blocked control is reported as blocked. This is the discipline the truthfulness track exists to score.

Reads the web with you

A real page opens on the shared stage, read-only, and the conversation grounds itself in what is actually rendered there — answering from the page, not from a memory of one.

Holds the floor while it works

Retrieval, tool calls and generation run behind a voice that stays present — no dead air, interruptions honored, recovery without losing the visitor's in-flight words. This is the Octra layer's floor discipline, the STO competence, and it is the track where orchestration is measured.

03On camera

Three moments from recorded sessions, reproduced from the research notes. The full set, with the surrounding argument, is in Larger Than Language §08.

Recorded sessionThe verification pair
Live UI panel header reading Playing
before — Playing:
Live UI panel header reading Paused
after — Paused:

“Pause it” → the caption reads “Pausing the video. Done, paused.” only as the interface chrome flips from Playing: to Paused:. The report is bound to the verified state change — the verification contract, on camera.

Recorded sessionPublished affordances
The Live UI panel showing an audio-only card with a Show the video control

A track retrieved in audio-only mode carries its own “Show the video” control — the surface publishing what can be done to it, which is what makes it voice-addressable at all.

Recorded sessionThe shared stage
The Live UI panel with a bar chart, a timeline and a metric row built during one answer

One answer's worth of interface: chart, timeline and metric row, composed while the sentence that explains them is still being spoken.

04Scored, in public

The families are evaluated on the Interactive Intelligence Bench (IIB‑1) — the benchmark we published for the competences nothing else measures: co-timed generation, on-screen reference, truthful verification, shared surface, floor discipline. Two composite readings are defined over its tracks, one per family, computed from the published per-track scores at the published weights:

Actra Score
38.7 / 100
LabsAI Studio · the Actra surface
T1 33.3 · T2 0 · T3 62.5 · T4 50 — weighted at the published IIB‑1 track weights. The action-side competences: arriving on time, resolving what was meant, telling the truth about what was done, reading the page together.
Octra Score
86.7 / 100
LabsAI Studio · the Octra layer
T5 86.7 — weighted at the published IIB‑1 track weights. Floor discipline - the Speech-to-Orchestration (STO) competence: keeping the conversation alive while the work happens.

The Actra Score aggregates the action-side tracks — co-timing, reference, verification, shared surface. The Octra Score is floor discipline, the STO competence. The scale is truthful by construction: the verification track scores a false success below zero, so a system that misreports what it did scores worse than one that cannot act at all. Every number is a recorded run with its rubric stated before the run and its transcript retained; the bench page carries the per-task detail, the capability matrix across ten systems, and the method with its limitations stated plainly.

05How the families learn

The learning discipline is the part of the program we hold most firmly, because it is the part that keeps the rest honest: weights are earned through the gate, not written into it.

Every session the surface serves leaves a trajectory — what was asked, what was staged, what was acted on, what verified, what failed — each one the record of a live AI Orch: the Studio's retrieval pipelines, recovery ladders, co-browse sessions and media rosters are executable coordination objects with participants, state and a lifecycle, written by hand in this generation, and exactly what a trained Octra learns to generate and adapt. That ledger is the family's native corpus: the interaction loop itself, recorded from the inside rather than scraped from the world. Candidate improvements — skills distilled from successful trajectories, persona overlays, policy adjustments — are evolved offline, replayed against a held-out set of real sessions, and admitted only on strict improvement through a gate in which execution metrics hold a veto and regression collapses the candidate back to the shipped behavior. Nothing self-modifies in production; a human promotes every candidate.

That is the first generation's learning loop, running today. The horizon it points at is stated in Larger Than Language §15: the ladder of hand-written decisions — reference resolution, routing, recovery, verification — is precisely the structure a trained Actra and a trained Octra internalize, and the trajectory corpus now accumulating is what they train against. The prediction target is already in production; the families exist to learn it.

06The LAIMA ladder

Within LAIMA — Labs AI Modeling Archetypes, the organizing framework for the Labsintelligence model registry — LaLaMo (Labs Large Models), LaMimo (Labs Medium Models) and Lasamo (Labs Small Models) are not further families. They are hierarchical forms — the three classes a Labs model ships as — and each family takes all of them: an Actra at LaLaMo class is the frontier form of the same family whose LaMimo form serves the working scale and whose Lasamo form runs light and close to the device. A general model may be trained broadly and then tuned to a surface; the grid below is the family × form map, with the shipped, scored system placed where it actually is and every unreleased rung marked as what it is.

Family LaLaMoLabs Large Models LaMimoLabs Medium Models LasamoLabs Small Models
Shipping todaythe system as composed LIVE
The first-generation Actra surface + Octra layer — LabsAI Studio with Live UI and Lens, scored 38.7 Actra / 86.7 Octra on IIB‑1. This row is a system, not a single model, and it is the specification the family rungs below are drawn from.
ActrasSpeech-to-Action IN RESEARCH & DEVELOPMENT
The full action vocabulary at frontier scale.
IN RESEARCH & DEVELOPMENT
The surface-tuned working scale.
IN RESEARCH & DEVELOPMENT
On-device and low-latency deployments.
OctrasSpeech-to-Orchestration IN RESEARCH & DEVELOPMENT
Full-fleet routing across the Labs intelligences.
IN RESEARCH & DEVELOPMENT
Session-scale orchestration.
IN RESEARCH & DEVELOPMENT
Embedded floor discipline.

Reading the grid plainly. One row is live and scored; six rungs are research. We mark them that way on purpose — a grid that fills itself in before the work is done has told you about the author, not the models — and the registry tense (“in research and development at Labsintelligence”) is the same tense the papers have always used for LAIMA.

07Epistemics

Claiming success on an action that did not land scores below zero. A plain “I tried that and it didn't work” scores above it.

The single scoring decision we defend hardest, and the one that costs us most. Confident false reports are the failure mode that destroys trust in an interface you are watching, and they are the one failure mode every system in this category shares. So the bench's verification track is signed, its adversarial items make false success reachable — a disabled control, a card removed before the command arrives — and the same discipline runs live: the surface's own reports are bound to verified state, a destructive command is held for confirmation, and “done” is a claim the interface has to earn. The companion note Knowing When to Speak carries the epistemics of the voice itself — presence, turn-taking, and when silence is the right answer.

08Where to meet it

The first-generation Actra surface and Octra layer are live in LabsAI Studio — speak with Chemí, and the interface builds while she talks. The research record behind this page: Knowing When to Speak (the voice loop), Larger Than Language (the model-class argument, the families in full), and the Interactive Intelligence Bench (the ruler, published before our own reflection in it). No API is offered today; teams who want to run IIB‑1 against their own systems are invited to — the submissions policy is on the bench page.

The argument this family exists to prove is stated in the companion note.

Read Larger Than Language

Contributors

Labsintelligence

Contributing authors: Maya E. Davis · Duránd F. Davis Jr.

Tell us what you think, join us

This research is published while the questions are still open, and the systems it describes are live. We would love for you to join us — and please share your thoughts at research@labsintelligence.ai.

Citation

Please cite this work as:

Davis, Maya E., and Davis, Duránd F., Jr., “Introducing Actras & Octras.” LabsAI Research Briefs, Labsintelligence — lab1 of Labs Companies, Inc., August 2026.

Or use the BibTeX citation:

@article{labsintelligence2026introducingactrasoctras,
  author  = {Davis, Maya E. and Davis, Duránd F., Jr.},
  title   = {Introducing Actras \& Octras},
  journal = {LabsAI Research Briefs},
  publisher = {Labsintelligence, lab1 of Labs Companies, Inc.},
  year    = {2026},
  month   = {august},
  url     = {https://labsintelligence.ai/research/labsai/introducing-actras-octras/},
}