Voice control of a digital audio workstation has never failed because speech recognition was not good enough. It has failed because a stateless command pipeline cannot be trusted with a session. This paper presents an architecture built to answer that failure — a session model, a bidirectional event bus, and three safety locks — with measurements from a reference implementation in continuous development since December 2024.
A stateless system does not know the session. Ask it to arm a track and it has no way to know whether that track exists, whether it is already armed, or whether arming it right now would disturb a take in progress. Only a live model of the session answers those questions; a command grammar cannot.
A stateless system cannot roll back safely. A compound command — create a bus, route two tracks through it, insert an effect — can fail partway. Without a record of what it already did, the system cannot undo the completed steps, and the user is left to repair the session by hand in software they may not know how to operate.
A stateless system is blind to the DAW's own actions. If the musician stops the transport with the mouse or deletes a track at the console, the voice layer keeps operating on a stale picture. The picture drifts further from reality with every change, and each command becomes a gamble.
Together these produce the thing that actually matters: a user who hits one failure does not trust the system enough to try again. In a recording environment, where a wrong move can destroy an unrepeatable take, trust is the entire product.
Full manuscript, 4–8 pages. Available on this page once the AES camera-ready version is filed.
An in-studio capture of a session run entirely by voice: commands issued from the instrument, and a compound operation rolled back cleanly after a deliberate failure.
System architecture overview, the routing waterfall, and latency distributions from the 1,000-command benchmark.
The measurements characterize the execution substrate: intent classification, command acknowledgment through the execution bridge, event propagation, and rollback success across three failure grades.
The paper does not claim task-level usability, and it reports no controlled user study. The reference implementation is one system tested against one commercial DAW. Generalizing the architecture to other hosts remains future work.
The reference implementation described in this paper is ALSE — Advanced Live Studio Environment, built by BenBlends.
For research, review, or integration conversations about the architecture:
This page carries the paper and its evidence. It is not a product page.