Building Voice OS
I started working on a project that spanned across five different repos, and the testing and debugging part was becoming hell. I was running five projects locally in the terminal, copying logs from all five into a Claude session to debug, and doing that over and over again. So I built crew, my open-source tool for workspace management, as a CLI that Claude can use. You register your projects and group them into a workspace, and each workspace can have multiple worktrees, which are just copies of the whole stack, so your own clones never get touched. Then I added install and run commands to each project, so Claude could start all the dev servers by itself, each in its own detached tmux session, read their logs while it worked and tested, and fix whatever broke.
That’s kind of when I first unlocked fully agentic work for me, because from that point on all I really cared about was setting up my workspaces correctly, and after that Claude did all the god’s work.
Why voice
So crew was already in a good place. I was using it for dev servers, it had its own plugin with skills, Claude was able to use it to read logs and it worked perfectly. But then I started thinking about what I’m actually doing the whole day, and it’s literally just switching between terminal windows and writing prompts, so I thought to myself, why couldn’t I do this with voice?
I’ve been working on voice agents for the past year now, so I have a lot of domain knowledge, I would say, especially around turn detection and building agents with tool calling. I mostly used realtime models, which have tool calls built in, although at work we actually ended up using the realtime model as just a glorified TTS. But I knew how to build a pipeline out of separate pieces, speech to text, text to speech and an LLM. I’d had this idea before and tried it, but I don’t think Claude or any other agent was good enough at that point, and I didn’t have the energy to dedicate much time to it.
Then one day I went for an hour-long walk and brainstormed this project with Claude using voice, because I also wanted to see how well I’d be able to communicate using just voice versus writing. So that whole hour went into what the agent would look like, and also the name. Voice OS is a name I first came up with at work. I had a vision for a fully programmable language tutor and wanted to call its architecture Voice OS internally, but the architecture changed and the name didn’t stick. With crew I thought it would be a good way to bring the name back, because you’re basically operating all your coding sessions with Claude, on your computer and remotely, using your voice, which I thought was pretty cool. So that’s how Voice OS came to be.
Soniox, Haiku, and why I stick with Claude
With all that domain knowledge and the brainstorming with Claude, we came up with a system where Soniox does both speech to text and text to speech. I like Soniox because they’re very accurate and very low latency, they work well across like 60 different languages, and they’re much, much cheaper than anybody else. The dev experience is also pretty good, it’s really easy to integrate and it just worked great.
For the LLM I chose Anthropic’s smaller model, Haiku, because it’s fast, and also if I’m building this for Claude I don’t see a good reason why I would use some other LLM.
I also love the Claude Code experience and how it handles skills and agents. I have a personal agent stack on top of Claude, agents with my philosophy on how to build software, how to write code and how to test code, and also my product philosophy. The whole of Voice OS was built using those, every single feature. It’s open source too, it’s in my repo called proxy, which is kind of a proxy of myself, I guess, even though that was also written by AI, so you know.
Building it by using it
After the walk I came up with a plan, asked Claude with my proxy agents to build me a prototype and went to bed. I used proxy solo, which is the one that does everything by itself, planning, implementing, reviewing and QA. When I woke up the prototype was built and it was working, which was pretty cool to see. I was able to switch between sessions with voice, start and restart dev servers and so on, so most of the functionality was there, but there were a lot of issues, especially with voice commands.
The way it’s built, there’s a kernel, which is just a Haiku prompt that gets what you said as text and decides whether it’s a command for Voice OS, like switching sessions, or something it should just forward to the session. It knows the context, it knows what you’re looking at right now, and based on all of that it tries to make the right decision about what to do next with what you said. This was the most problematic part. But I think one good decision I made from the very beginning was to start with evals, so we have evals for everything, tests for everything and end-to-end tests for everything, which makes it really easy to iterate.
Markdownflowchart
```mermaid flowchart LR you["me, talking"] --> stt["Soniox<br/>speech to text"] stt --> kernel["kernel<br/>Haiku"]:::key kernel -->|"switch, approve, note"| tool["Voice OS<br/>does it"] kernel -->|"words for the work"| session["Claude session<br/>my exact words"] session --> tts["Soniox<br/>text to speech"] tool --> tts classDef key fill:#fff,color:#000 ```
The other thing I did very early on was add a debug note command. When I was testing and something was wrong, I could ask it to add a debug note about it. The note knows the context around it, it has a timestamp and it can look up the logs around that moment, what I was saying and what the replies were. So for example when I said “can you switch to that session” and it did something else, I just added a debug note saying this is what happened, blah blah blah. Then I used my voice again to go to the session that works on crew itself and asked it to read the debug notes and fix the issues using the evals. I was basically building it by using it, adding debug notes and feature requests and asking the crew session to build them, and that was kind of the building loop.
Markdownflowchart
```mermaid flowchart LR use["using Voice OS<br/>by voice"] --> note["debug note<br/>time, logs, what was said"]:::key note --> session["crew session<br/>reads the notes"] session --> evals["fix against<br/>the evals"] evals --> push["new build"] push --> use classDef key fill:#fff,color:#000 ```
Remote machines
I built crew to make my time at work easier, and I’ve used it for work since the start. Whenever I saw something missing or something that could be better I just updated it, and I always had a Claude instance open for crew. One of the issues was that I don’t only work on this Mac. My old work laptop is also a Mac I can build on, and I can leave it with a task while I’m commuting with this one closed. I also have a VM for personal use. That’s when I got the idea for what I think is now one of the main features, remote machines.
You install crew on multiple devices. On the remote ones you start crew as a remote, and on your main device you start it normally and add the others with their SSH host. Crew then connects to those machines over SSH and everything you do flows to the sessions there. So I can control Claude instances on other machines, it’s all managed by crew and always online, and I can work on five different remotes at the same time, switching between them by voice effortlessly.
Markdownflowchart
```mermaid flowchart LR mac["this Mac<br/>main: the page, voice, keys"]:::key mac --> local["sessions<br/>on this Mac"] mac -->|ssh| laptop["old work laptop<br/>sessions, dev servers"] mac -->|ssh| vm["personal VM<br/>sessions, dev servers"] classDef key fill:#fff,color:#000 ```
When I was developing it, it was kind of annoying that every time I did a release I had to log into one remote and update it, then update the other remote, then update the main, so now the main distributes updates. When the main installs its new binary it gets the remotes to do it as well, because all the remotes have to be on the same version as the main. And at some point I didn’t want to go through the release process every time I made a change just to test it. It’s only me using it, so why would I overcomplicate things? Now the main can also push a dev build: it installs it on the remotes first, then on itself, reloads, and everything connects again.
MarkdownsequenceDiagram
```mermaid sequenceDiagram participant M as this Mac (main) participant L as work laptop participant V as personal VM M->>L: install the new build M->>V: install the new build L-->>M: updated V-->>M: updated M->>M: install itself, reload M->>L: reconnect M->>V: reconnect ```
This is one of the features I use the most, I always have a few sessions from each remote active, and I think it kind of solved remote work for me. Before, I just had ten terminal windows open, I was losing SSH connections, I had to log in again, reopen, find the history, open the right chat, blah blah blah. Now crew does all of that for me and the chat just continues. I don’t know how good one never-ending conversation is for the token economy, I guess, and that’s something I might have to work on depending on how my usage goes. But I’ve never hit the limits, even running two different work projects plus crew and my proxy agents at the same time.
The redesign, and why crew moved to the web
Then I started a redesign, which began with a vision of this dramatic fade-in animation. When I opened Voice OS I wanted it to look cool, like a full-screen overlay transitioning into crew, so that’s what I did first. And when I really liked the animation I was like, okay, but fuck, now it doesn’t really fit anything else. Crew was a TUI and Voice OS was web, so when I built Voice OS my stack kind of shifted, because before that it was all in the terminal. It was either the CLI, direct commands with a lot of examples and help, which agents worked really well with, or the TUI, where you used your keyboard to go through it and pick what you wanted to do, which was more of a UI but still text-based.
So with the redesign I thought it would be really nice if, when you open crew, you get Voice OS and also Set up, which is basically everything crew was able to do in the TUI, but on the web. That way I have one place for it all. A TUI is okay, and when it gets too complex it still works, it’s just not manageable anymore. I used an Apple design skill for the redesign, because I really like how they do UX and I kind of like the clean Apple look.
The redesign was also kind of a night job. I first asked Claude to draft the Voice OS redesign in artifacts, and I loved how it looked, so I had it draft the whole TUI as web pages as well. I reviewed it and then just told Claude to build it. Before that, crew was the front and Voice OS was kind of a side server you could download and start. Because everything moved to the web, the web became the main thing, so now when you type crew it opens the web page. The only TUI thing I left is crew launch, which is for when you want a normal Claude Code session in the terminal for your workspace.
If you want to try it
Crew runs on macOS and Linux, with Claude Code as the agent. Install it and run crew, and it opens its page in your browser, finds the repos you already have and walks you through making your first workspace. Voice OS needs two API keys, Anthropic for the kernel and Soniox for speech.
$ curl -fsSL https://raw.githubusercontent.com/FurlanLuka/crew/main/install.sh | sh
$ crewIt’s all on GitHub and at getcrew.sh, and the release notes say what changed in every version. I build crew with crew every day, so it moves fast.
What’s next
I’m going to keep updating it and adding features as I need them, the same way I have so far, because I use it every day, for work and for my own projects. If you have feature ideas or you find bugs, everyone is welcome to contribute, so open an issue or a pull request on GitHub.