LabsAI · Research Brief · Model Classes
Speech-to-Action, and Speech-to-Orchestration.
August 16, 2026
Speech-to-text made machines hear. Text-to-speech made them speak. Neither made them do. STA Models — Speech-to-Action — and STO Models — Speech-to-Orchestration — are the two classes a spoken system needs the moment it stops merely talking and starts acting on a live interface: one converts spoken intent into action on the surface, the other converts it into the coordination of the intelligences doing the work.
An STA Model converts spoken intent, conversation, application state and the vocabulary of available actions into action on a live surface — co-timed with the voice that explains it. The competences that define the class are the ones a visitor actually feels: a referent resolving to the right on-screen object (“pause it”, “the second one”, “no — the other one”); a clarifying question on a genuinely ambiguous reference instead of a guess; and a report bound to a verified outcome — done meaning checked, not assumed. Actras are the Labs family implementing STA; in generative terms, an Actra is a GAT, a Generative Action Transducer.
An STO Model converts spoken intent into coordination: which model or tool handles the work, what runs in parallel, when to retrieve, when to escalate, how to recover — all of it steered by live speech at conversational tempo, while a human holds the floor. Its output is orchestration policy, and the policy runs as an AI Orch: the executable coordination object with participants, dependencies, state and a lifecycle. Octras are the Labs family implementing STO; in generative terms, an Octra is a GOT, a Generative Orchestration Transducer.
STA decides what the surface does. STO decides which intelligence does it.
Throughout the platform, this work arrives spoken: a visitor’s sentence lands as Speech‑to‑Action Voice Commands — the unit of spoken work an STA Model interprets and executes against the surface — and as Speech‑to‑Orchestration Voice Commands, the unit of spoken coordination an STO Model runs as policy. One sentence may carry several of each; the classes decide where each lands. Beneath the classes sit their programming interfaces — IPI exposes what the surface can do; OPI exposes how intelligence can be coordinated.
The classes are stated in full in Larger Than Language; the voice loop that runs them is specified in Knowing When to Speak; and both are scored in public on the Interactive Intelligence Bench — the Actra Score over the action-side tracks, the Octra Score over floor discipline, with a truthfulness penalty that scores a false “done” below an honest failure. The first-generation STA surface and STO layer are live in LabsAI Studio.
The families implementing both classes are presented in the companion announcement.
Read Introducing Actras & Octras →Contributing authors: Maya E. Davis · Duránd F. Davis Jr.
This research is published while the questions are still open, and the systems it describes are live. We would love for you to join us — and please share your thoughts at research@labsintelligence.ai.
Please cite this work as:
Davis, Maya E., and Davis, Duránd F., Jr., “Introducing STA & STO.” LabsAI Research Briefs, Labsintelligence — lab1 of Labs Companies, Inc., August 2026.
Or use the BibTeX citation:
@article{labsintelligence2026introducingstasto,
author = {Davis, Maya E. and Davis, Duránd F., Jr.},
title = {Introducing STA \& STO},
journal = {LabsAI Research Briefs},
publisher = {Labsintelligence, lab1 of Labs Companies, Inc.},
year = {2026},
month = {august},
url = {https://labsintelligence.ai/research/labsai/introducing-sta-sto/},
}