Jev in production › Agent tool and action selection
Voice-driven macOS computer use where Jev picks the next operation and target from the Accessibility tree, with no screenshots, and the app executes and re-reads the screen. About 0.3–1.5 s per step (author).
Voice and typed computer use for macOS. You say what you want. Jev picks the next on-screen action. macOS performs it. No screenshots: the app reads the screen through the Accessibility tree.
In setup: save your TypeSafe API key (stored in the Keychain), then allow Accessibility, microphone and speech.
Hold Control–Option–Space, speak, release. Change this combination in Settings → Choose your voice shortcut → Change shortcut; press the key and any modifiers you want. Your choice is saved across launches. Escape cancels recording, and a conflicting shortcut leaves the previous choice in place. macOS-reserved combinations and modifier-only shortcuts are not supported. Escape cancels. Or type a command in the widget, or from a shell:...
For the project's own README, linking back here: