LabsAI · Research Brief · Model Families
Generative Action, Orchestration and Visualization Transducers.
August 16, 2026 · Updated September 2026
One acronym made the last decade of AI legible: GPT, the generative pre-trained transformer, named the move from analyzing language to generating it. The same move is now happening three times over, and each deserves its letters. A GAT — a Generative Action Transducer — generates action. A GOT — a Generative Orchestration Transducer — generates coordination. A GVT — a Generative Visualization Transducer — generates visualization. This page defines all three, and names the Labs families that implement them.
A GAT takes speech, conversation, application state, memory, intent and the vocabulary of available actions, and generates action: sequences, parameters, state transitions, tool operations, confirmations — co-timed with what the voice says. Spoken, this class is Speech‑to‑Action (STA). Actras — Action-Centered Transducers for Reasoning and Agentic Systems — are the Labs GAT family, and Actras are to speech what VLA models are to vision: the vision-language-action class made the same move for embodiment, emitting actions as tokens in the same stream as perception.
A GOT takes state, objective and resources, and generates coordination: model and tool selection, decomposition, parallel versus sequential execution, escalation, verification paths, termination. Spoken, this class is Speech‑to‑Orchestration (STO). Octras — Orchestration-Centered Transducers for Routing and Agentic Systems — are the Labs GOT family. What a GOT generates runs as an AI Orch: the executable coordination object that binds models, agents, tools, workflows, memory, policies, data and compute into one live operation — Orch Models are intelligence; AI Orchs are execution structures.
A GVT takes speech, conversation, intent, world state, application state, available information, visual context and representational constraints, and generates visualization: what should become visible, which visual representation best expresses an intent or state, how information is spatially composed, when visual elements appear, how they evolve, and how the resulting visual environment remains synchronized with speech, action and orchestration — generating the visual representation itself rather than merely generating words describing it. Spoken, this class is Speech‑to‑Visualization (STV). Vistras — Visualization-Centered Transducers for Representation and Agentic Systems — are the Labs GVT family. A GVT is not text‑to‑image with speech transcription placed in front of it, and it is not voice‑controlled generative UI; those can be components of an STV system, and neither defines the class. The distinction is the prediction target: what should this become visually?
A transducer is named for the conversion it performs, not the topology that performs it. A GAT converts speech and world state into action; a GOT converts objective and resources into policy; a GVT converts speech, intent, context and state into visual representation. That keeps all three categories architecture-agnostic on purpose — the families may span decoder-only stacks, state-space models, diffusion decision models, RL policies and neuro-symbolic hybrids — while the transformer lineage, from Attention Is All You Need through GPT, is cited as ancestry rather than claimed as a constraint. In that precise sense, Transducers are post-transformer as a framework, not as a topology: the class is named for the conversion it performs, and the transformer is ancestry, not definition.
What GPT named for generative language, GAT names for generative action, GOT for generated coordination and GVT for generated visualization.
Spoken, the classes’ work arrives as voice commands — Speech‑to‑Action Voice Commands for what a GAT generates, Speech‑to‑Orchestration Voice Commands for what a GOT generates, Speech‑to‑Visualization Voice Commands for what a GVT generates — and the execution-side classes are programmed against their own interfaces: the IPI on the action side, the OPI on the orchestration side.
The action and orchestration categories are scored in public on the Interactive Intelligence Bench (IIB‑1) — the Actra Score over the action-side tracks, the Octra Score over floor discipline, every number a recorded run. A dedicated Vistra evaluation framework should measure visualization competence independently rather than retroactively assigning STV performance from Actra or Octra scores. The full argument lives in Larger Than Language, and the first-generation Actra surface and Octra layer are live in LabsAI Studio.
The families themselves, with the LAIMA ladder they ship across, are presented in the companion announcements.
Read Introducing Actras, Octras & Vistras → Read Introducing Speech-to-Visualization →Contributing authors: Maya E. Davis · Duránd F. Davis Jr.
This research is published while the questions are still open, and the systems it describes are live. We would love for you to join us — and please share your thoughts at research@labsintelligence.ai.
Please cite this work as:
Davis, Maya E., and Davis, Duránd F., Jr., “Introducing GAT, GOT & GVT.” LabsAI Research Briefs, Labsintelligence — lab1 of Labs Companies, Inc., August 2026; expanded September 2026.
Or use the BibTeX citation:
@article{labsintelligence2026introducinggatgot,
author = {Davis, Maya E. and Davis, Duránd F., Jr.},
title = {Introducing GAT, GOT \& GVT},
journal = {LabsAI Research Briefs},
publisher = {Labsintelligence, lab1 of Labs Companies, Inc.},
year = {2026},
month = {august},
note = {Expanded September 2026},
url = {https://labsintelligence.ai/research/labsai/introducing-gat-got/},
}